AWS云工程师IV
AWS Cloud Engineer IV
职位概述
L4云DevOps工程师是云架构和运维工程的技术专家(SME),专注于AWS、Terraform(IaC)、操作系统、Kubernetes(EKS)和现代DevOps实践。该职位设计、实施和支持云基础设施和部署自动化,主导事件响应和灾难恢复活动,指导初级工程师,并与相关方协作,交付符合业务目标的安全、可靠和可扩展的解决方案。
职业级别概述
- 需要具备在AWS、Terraform、操作系统管理、Kubernetes和DevOps工具方面的深厚技术专长。
- 带领他人解决复杂的技术问题,并在项目中提供技术领导力。
- 独立处理复杂的工程任务;仅在最复杂、模糊的情况下进行升级或寻求指导。
- 可能为L1/L2工程师提供职能领导和指导。
关键能力
- 强大的沟通和利益相关者管理能力;能够将技术概念传达给非技术人员。
- 影响力和辅导:指导团队成员,推广最佳实践,推动持续改进。
- 分析性问题解决:诊断复杂的生产问题并识别根本原因和长期解决方案。
- 协作:跨学科合作设计、实施和运营云解决方案。
- 项目和时间管理:优先处理工作,跟踪交付成果,并在平衡多个项目的同时按时完成任务。
- 技术文档:创建清晰的操作手册、标准操作流程、架构图和标准。
主要职责
- 作为复杂AWS、操作系统、Kubernetes、基础设施和部署问题的L3升级点和技术专家。
- 使用Terraform(以及可选的CloudFormation)设计和实现基础设施即代码,构建可重用模块并执行编码标准。
- 实施和维护CI/CD流水线和GitOps实践(ArgoCD),包括流水线设计、环境发布、回滚和漂移修复。
- 使用Ansible/AWX、脚本(Bash、PowerShell、Python)和其他自动化工具开发配置管理和操作任务的自动化。
- 设计、部署和维护Kubernetes工作负载(Amazon EKS)、Helm图表和特定环境的模板。
- 执行Linux和Windows管理:加固、补丁、性能调优和自动化部署。
- 主导
查看英文原文
Job Profile Summary
The L4 Cloud DevOps Engineer serves as a technical subject matter expert (SME) for cloud architecture and operational engineering, focusing on AWS, Terraform (IaC), operating systems, Kubernetes (EKS), and modern DevOps practices. The role designs, implements, and supports cloud infrastructure and deployment automation, leads incident response and DR activities, mentors junior engineers, and collaborates with stakeholders to deliver secure, reliable, and scalable solutions that align with business objectives.
Career Level Summary
- Requires deep technical expertise across AWS, Terraform, OS administration, Kubernetes, and DevOps tooling.
- Leads others to solve complex technical problems and provides technical leadership on projects.
- Works independently on sophisticated engineering tasks; escalates or seeks guidance for only the most complex, ambiguous situations.
- May provide functional leadership and mentorship to L1/L2 engineers.
Critical Competencies
- Strong communication and stakeholder management; able to translate technical concepts for non-technical audiences.
- Influence and coaching: mentor team members, promote best practices, and drive continuous improvement.
- Analytical problem solving: diagnosing complex production issues and identifying root cause and long-term fixes.
- Collaboration: work across disciplines to design, implement, and operate cloud solutions.
- Project and time management: prioritize work, track deliverables, and meet deadlines while balancing multiple initiatives.
- Technical documentation: create clear runbooks, SOPs, architecture diagrams, and standards.
Key Responsibilities
- Act as the L3 escalation point and technical SME for complex AWS, OS, Kubernetes, infrastructure and deployment issues.
- Design and implement infrastructure as code using Terraform (and optionally CloudFormation), building reusable modules and enforcing coding standards.
- Implement and maintain CI/CD pipelines and GitOps practices (ArgoCD), including pipeline design, environment promotion, rollback, and drift remediation.
- Develop automation for configuration management and operational tasks using Ansible/AWX, scripting (Bash, PowerShell, Python), and other automation tools.
- Design, deploy and maintain Kubernetes workloads (Amazon EKS), Helm charts, and environment-specific templating.
- Perform Linux and Windows administration: hardening, patching, performance tuning, and automated deployments.
- Lead incident response and major production incidents; perform RCA and implement permanent automated fixes.
- Lead and coordinate Disaster Recovery planning and testing, including runbooks, failover/failback, and validation.
- Define and enforce technical standards for IaC, Kubernetes, CI/CD, monitoring, logging, and operational processes.
- Improve observability: monitoring, alerting, logging and operational reliability across platforms.
- Mentor and provide technical guidance to L1 and L2 engineers; review and approve infrastructure and deployment changes.
- Create and maintain SOPs, runbooks, architecture documentation, and DR documentation.
- Collaborate with product and customer engineering teams to deliver cloud-based solutions, migrations, and optimizations.
Experience
10+ years of relevant experience in cloud engineering, infrastructure, DevOps, or a related field is required.
Knowledge
- Expert-level knowledge of AWS services, architecture patterns, security, networking, and operational best practices.
- Deep understanding of Infrastructure as Code concepts and Terraform module design, state management and security.
- Extensive knowledge of Linux and Windows internals, administration, automation, and security hardening.
- In-depth Kubernetes architecture and operations experience (preferably Amazon EKS), including Helm and GitOps patterns.
- Familiarity with common security and compliance frameworks (e.g., NIST, HIPAA, PCI) and secure operational controls.
Skills
The candidate must demonstrate hands-on, expert-level skills across the following areas:
- Cloud & IaC: Expert in AWS and Terraform (including Terragrunt patterns), reusable module development, state management, and IaC security.
- Kubernetes & Containers: Design, operate and troubleshoot EKS clusters, Helm chart development, container best practices, and GitOps (ArgoCD).
- CI/CD & Automation: Build and maintain CI/CD pipelines, artifact management, automated testing, and deployment automation (Jenkins, GitHub Actions, GitLab CI, etc.).
- Configuration Management: Use Ansible/AWX for system configuration, patching, application deployment and repetitive operational tasks.
- Operating Systems: Expert-level Linux administration and solid Windows server administration; automation via scripting (Bash, PowerShell, Python).
- Version Control: Strong Git skills (branching strategies, code review, hooks, access controls).
- Monitoring & Observability: Implementing monitoring, logging, and alerting (CloudWatch, Prometheus, ELK/EFK or similar).
- Security & IAM: Implement secure IAM practices, secret management, network security and encryption in cloud environments.
- Disaster Recovery & Resilience: DR planning, regular testing, and runbook creation for reliable failover and recovery.
L4 / SME Expectations
- Lead and own complex engineering changes and transformations across cloud platforms.
- Drive automation, standardization, and reliability improvements in operations and delivery.
- Define technical standards, review peer code, and enforce best engineering practices.
- Act as a technical escalation point for production incidents and lead root cause analysis.
- Provide mentorship and hands-on guidance to junior engineers, elevating team capability.
Certifications
- Preferred: AWS Certifications (Solutions Architect, DevOps Engineer Professional).
- Beneficial: CNCF/Kubernetes Certifications (CKA, CKAD, CKS), Azure/GCP certifications as applicable.
About Rackspace Technology
We are the multicloud solutions experts. We combine our expertise with the world’s leading technologies — across applications, data and security — to deliver end-to-end solutions. We have a proven record of advising customers based on their business challenges, designing solutions that scale, building and managing those solutions, and optimizing returns into the future. Named a best place to work, year after year according to Fortune, Forbes and Glassdoor, we attract and develop world-class talent. Join us on our mission to embrace technology, empower customers and deliver the future.
More on Rackspace Technology
Though we’re all different, Rackers thrive through our connection to a central goal: to be a valued member of a winning team on an inspiring mission. We bring our whole selves to work every day. And we embrace the notion that unique perspectives fuel innovation and enable us to best serve our customers and communities around the globe. We welcome you to apply today and want you to know that we are committed to offering equal employment opportunity without regard to age, color, disability, gender reassignment or identity or expression, genetic information, marital or civil partner status, pregnancy or maternity status, military or veteran status, nationality, ethnic or national origin, race, religion or belief, sexual orientation, or any legally protected characteristic. If you have a disability or special need that requires accommodation, please let us know.
Originally posted on Himalayas