高级DevOps工程师
Senior DevOps Engineer
我们正在寻找一位经验丰富的高级DevOps工程师,负责维护和扩展我们的多云基础设施,提供值班支持,并通过结对编程与工程团队紧密合作。你将为确保客户的系统可用性、稳定性和可扩展性做出关键贡献。
此职位最适合一位能够接受部分所有权,并在低工时的值班工作中提供高影响力支持的资深工程师。
关键关注领域
- 多云基础设施(AWS为主;Azure和GCP)
- 容器编排(Docker、Docker Swarm;熟悉Podman或Kubernetes者优先)
- 客户连接网络(VPC、VPN、安全组、子网)
- 快速事件响应和故障排查
- 结对编程和知识共享
主要职责
- 为生产事件和基础设施问题提供值班支持;快速排查并解决
- 与工程师结对编程以排查问题并分享基础设施最佳实践
- 维护和排查云基础设施,包括Docker容器和Docker Swarm编排
- 管理网络和客户连接模式
- 维护Terraform/Terragrunt配置并自动化部署流程
- 监控系统健康和性能,并实施告警优化
- 实施安全最佳实践,管理IAM角色并配置密钥
- 记录基础设施架构、操作手册和操作流程
- 协作开发CI/CD流程(主要使用GitHub Actions)并支持自动化项目
所需技能与经验
- 3年以上AWS生产经验(VPC、EC2、S3、IAM、CloudWatch);熟悉Azure或GCP
- 具有Docker和Docker Swarm的生产经验,包括容器网络和服务发现
- 熟练使用Terraform;熟悉Terragrunt者优先
- 具有GitHub Actions和脚本(Bash、Python)的CI/CD经验
- 具有参与值班轮班和事件响应的经验
- 良好的沟通能力,熟悉协作式结对编程
优选技能
- 使用Python和/或Bash进行基础设施自动化
- 多云架构模式
- 数据流水线基础设施经验
- 安全最佳实践和成本优化经验
- 强大的技术文档编写能力
技术栈
- 云:AWS(主要),Azure,GCP
- 容器:Docker,Docker Swarm,Pod
查看英文原文
Overview
We’re seeking an experienced Senior DevOps Engineer to maintain and expand our multi-cloud infrastructure, provide on-call support, and collaborate closely with our engineering team through pair programming. You will be a key contributor to ensuring uptime, stability, and scalability for our clients’ systems.
This role is best suited for a senior engineer who is comfortable with fractional ownership and delivering high-impact support in a low-hour, on-call engagement.
Key Focus Areas
- Multi-cloud infrastructure (AWS primary; Azure and GCP)
- Container orchestration (Docker, Docker Swarm; familiarity with Podman or Kubernetes a plus)
- Networking for client connectivity (VPCs, VPNs, security groups, subnets)
- Rapid incident response and troubleshooting
- Collaborative pair programming and knowledge sharing
Key Responsibilities
- Provide on-call support for production incidents and infrastructure issues; troubleshoot and resolve rapidly
- Pair program with engineers to troubleshoot issues and share infrastructure best practices
- Maintain and troubleshoot cloud infrastructure, including Docker containers and Docker Swarm orchestration
- Manage networking and client connectivity patterns
- Maintain Terraform/Terragrunt configurations and automate deployment processes
- Monitor system health and performance, and implement alerting improvements
- Implement security best practices, manage IAM roles, and configure secrets
- Document infrastructure architecture, runbooks, and operational procedures
- Collaborate on CI/CD workflows (primarily GitHub Actions) and support automation initiatives
Required Skills & Experience
- 3+ years of production experience with AWS (VPC, EC2, S3, IAM, CloudWatch); familiarity with Azure or GCP
- Production experience with Docker and Docker Swarm, including container networking and service discovery
- Strong Terraform experience; Terragrunt preferred
- CI/CD experience with GitHub Actions and scripting (Bash, Python)
- Experience participating in on-call rotations and incident response
- Strong communication skills and comfort with collaborative pair programming
Preferred Skills
- Python and/or Bash for infrastructure automation
- Multi-cloud architecture patterns
- Data pipeline infrastructure experience
- Security best practices and cost optimization experience
- Strong technical documentation skills
Technical Stack
- Cloud: AWS (primary), Azure, GCP
- Containers: Docker, Docker Swarm, Podman, Kubernetes
- IaC: Terraform, Terragrunt
- CI/CD: GitHub Actions
- Scripting: Python, Bash
- Monitoring: CloudWatch, Azure Monitor
- Storage: S3, ADLS Gen2
- Secrets: AWS Secrets Manager, Azure Key Vault
Engagement Details
- Approximately 5–15 hours per month
- Contractor/Consultant role, fully remote
- On-call availability required; flexible schedule
- Long-term engagement with opportunity for expanded scope and responsibility
Originally posted on Himalayas