DevOps负责人
DevOps Lead
职位概述
DevOps 主任工程师负责在多云环境(AWS、Azure 或 GCP)中领导可扩展、安全、高弹性的云基础设施和 CI/CD 流水线的设计、实施和优化。该角色推动基础设施即代码的采用、容器编排和可观测性实践,同时确保系统的高可靠性与运营效率。DevOps 主任工程师指导和培养 DevOps 团队成员,与工程、安全和产品团队紧密协作,并将自动化、治理和成本优化的最佳实践融入其中,以支持大规模创新和平台稳定性。
主要职责
在 AWS、Azure 或 GCP 环境中领导可扩展、高弹性的云基础设施的设计和实施
使用 Jenkins、GitLab CI、GitHub Actions 或 Azure DevOps 等工具构建和优化 CI/CD 流水线
使用 Terraform、Ansible 或类似自动化工具推广基础设施即代码实践
使用 Docker 设计和管理容器化环境,并通过 Kubernetes 或托管 Kubernetes 服务编排工作负载
使用 Prometheus、Grafana、Datadog 或云原生监控解决方案建立并增强监控、日志和可观测性平台
为 DevOps 团队成员提供技术指导、辅导和支持
与工程、安全和产品团队跨职能协作,优化发布周期并提高部署可靠性
实施并执行云安全最佳实践、治理标准和合规要求
推动云成本优化策略和基础设施效率改进计划
在平台和工程团队中推广自动化、可靠性和持续改进的文化
排查复杂的基础设施和部署问题,确保对业务运营的最小干扰
参与文档编写、标准制定和长期平台架构战略
其他指派的任务
任职要求
- 计算机科学、信息技术或相关领域的学士学位(或同等经验)
- 8 年以上 DevOps、SRE 或平台工程经验
- 至少 2 年在领导或团队负责人岗位的经验
- 在多云环境(AWS、Azure 或 GCP)中有扎实的经验
- 在 CI/CD 工具链方面有深入经验(Jenkins、GitLab CI、GitHub Actions、Azure DevOps)
查看英文原文
Position Summary
The DevOps Lead Engineer is responsible for leading the design, implementation, and optimization of scalable, secure, and resilient cloud infrastructure and CI/CD pipelines across multi-cloud environments (AWS, Azure, or GCP). This role drives infrastructure-as-code adoption, container orchestration, and observability practices while ensuring high system reliability and operational efficiency. The DevOps Lead Engineer mentors and guides DevOps team members, collaborates closely with engineering, security, and product teams, and embeds best practices in automation, governance, and cost optimization to support innovation and platform stability at scale.
Key Responsibilities
Lead the design and implementation of scalable, resilient cloud infrastructure across AWS, Azure, or GCP environments
Architect, build, and optimize CI/CD pipelines using tools such as Jenkins, GitLab CI, GitHub Actions, or Azure DevOps
Champion infrastructure-as-code practices using Terraform, Ansible, or similar automation tools
Design and manage containerized environments using Docker and orchestrate workloads with Kubernetes or managed Kubernetes services
Establish and enhance monitoring, logging, and observability platforms using tools such as Prometheus, Grafana, Datadog, or cloud-native monitoring solutions
Lead DevOps team members by providing technical guidance, mentorship, and performance support
Collaborate cross-functionally with engineering, security, and product teams to streamline release cycles and improve deployment reliability
Implement and enforce cloud security best practices, governance standards, and compliance requirements
Drive cloud cost optimization strategies and infrastructure efficiency initiatives
Promote a culture of automation, reliability, and continuous improvement across platform and engineering teams
Troubleshoot complex infrastructure and deployment issues, ensuring minimal disruption to business operations
Contribute to documentation, standards development, and long-term platform architecture strategy
Other duties as assigned
Qualifications
- Bachelor’s degree in Computer Science, Information Technology, or related field (or equivalent experience)
- 8+ years of experience in DevOps, SRE, or platform engineering
- Minimum 2 years in a leadership or team-lead capacity
- Strong expertise in multi-cloud environments (AWS, Azure, or GCP)
- Deep experience with CI/CD tooling (Jenkins, GitLab CI, GitHub Actions, Azure DevOps)
- Proficiency in infrastructure-as-code tools (Terraform, Ansible)
- Hands-on experience with containerization (Docker) and Kubernetes orchestration
- Experience with monitoring and observability platforms (Prometheus, Grafana, Datadog)
- Strong scripting skills (Python, Bash, or similar)
- Solid understanding of networking, cloud security best practices, and cost optimization strategies
- Proven ability to lead technical teams and collaborate cross-functionally
- Relevant cloud certifications preferred
- Alignment with RTS Core Values
Originally posted on Himalayas