基础设施云工程师
Infrastructure Cloud Engineer
职位概述:
Creative Chaos 正在寻找一位实战型云工程师,负责在 Azure 和 AWS 上设计、自动化、保障安全并运营云工作负载。该职位负责核心平台组件,包括基础设施即代码(Terraform)、Kubernetes(AKS/EKS)、安全网络、CI/CD 支持、可观测性以及 FinOps。您将与 DevOps、软件和网页工程团队紧密合作,交付高弹性、可扩展且合规的云平台。理想的候选人应具备多云架构、Kubernetes 运维、身份与访问管理、安全策略、自动化及平台可靠性方面的扎实能力,并以务实、自动化优先的思维推动云工程。
主要职责:
平台工程
- 在 Azure 和 AWS 上设计并实现落地区(中心辐射型、策略防护)。
- 构建并维护 Terraform 模块、工作区、远程状态以及自动化环境部署(从开发到生产)。
- 运维并加固 AKS/EKS 集群,包括节点池、自动扩展、入口、镜像扫描/签名以及零停机升级。
- 实现并增强 CI/CD 流水线(GitHub Actions、Azure DevOps、Jenkins),用于构建、测试、扫描、部署和受控发布。
- 支持应用平台,如 API 管理/API 网关、Azure Functions/AWS Lambda 以及消息服务(Service Bus、SNS/SQS、EventBridge)。
- 负责 Azure Monitor、Log Analytics、App Insights、CloudWatch、X-Ray 和 OpenTelemetry 的可观测性,确保可操作的警报、运行手册、SLI/SLO 以及值班参与。
- 推动 FinOps 实践,包括标签标准、成本分配、资源优化、预留实例/节省计划、出站流量优化以及 Well-Architected 审查。
安全、治理与运维
- 接入日志/遥测数据,并与 SIEM 集成数据源。
- 使用 Azure Policy、AWS Config、Defender for Cloud、Security Hub、GuardDuty 和 WAF 策略实现并维护安全策略。
- 在 Entra ID(PIM、托管标识)和 AWS IAM/Identity Center 中实施最小权限访问,包括 CI/CD 的工作负载身份联合。
- 通过 IaC 优先的工作流管理变更控制和审计流程,同时维护运行手册和架构决策记录。
- 维护 Kubernetes、节点操作系统/AMIs、容器镜像和托管服务的补丁和版本卫生,包括自动化漂移检测。
- 领导跨 Azure/AWS 的事件调查,执行根本原因分析并实施预防性控制。
查看英文原文
Job Summary:
Creative Chaos is seeking a hands-on Cloud Engineer to design, automate, secure, and operate cloud workloads across Azure and AWS. This role owns core platform components including infrastructure as code (Terraform), Kubernetes (AKS/EKS), secure networking, CI/CD enablement, observability, and FinOps. You will work closely with DevOps, software, and web engineering teams to deliver resilient, scalable, and compliant cloud platforms. The ideal candidate is strong in multi-cloud architecture, Kubernetes operations, identity and access management, security guardrails, automation, and platform reliability—bringing a pragmatic, automation-first mindset to cloud engineering.
Key Responsibilities:
Platform Engineering
- Design and implement landing zones (hub-and-spoke, policy guardrails) across Azure and AWS.
- Build and maintain Terraform modules, workspaces, remote state, and automated environment provisioning (dev → prod).
- Operate and harden AKS/EKS clusters including node pools, autoscaling, ingress, image scanning/signing, and zero-downtime upgrades.
- Implement and enhance CI/CD pipelines (GitHub Actions, Azure DevOps, Jenkins) for build, test, scan, deploy, and gated promotions.
- Enable application platforms such as API Management/API Gateway, Azure Functions/AWS Lambda, and messaging services (Service Bus, SNS/SQS, EventBridge).
- Own observability across Azure Monitor, Log Analytics, App Insights, CloudWatch, X-Ray, and OpenTelemetry, ensuring actionable alerts, runbooks, SLIs/SLOs, and on-call participation.
- Drive FinOps practices including tagging standards, cost allocation, rightsizing, reserved instances/savings plans, egress optimization, and Well-Architected reviews.
Security, Governance & Operations
- Onboard logs/telemetry and integrate data sources with the SIEM.
- Implement and maintain security guardrails using Azure Policy, AWS Config, Defender for Cloud, Security Hub, GuardDuty, and WAF policies.
- Enforce least-privilege access across Entra ID (PIM, managed identities) and AWS IAM/Identity Center, including workload identity federation for CI/CD.
- Manage change control and audit processes through IaC-first workflows, along with runbooks and architectural decision records.
- Maintain patch and version hygiene for Kubernetes, node OS/AMIs, container images, and managed services, including automated drift detection.
- Lead incident investigations across Azure/AWS, perform RCA, and implement preventative controls (policies, guardrails, pipeline checks).
- Provide architectural input on security, reliability, networking, and cost during design reviews.
Requirements
- Bachelors in IT, CS or related field
- Minimum 5 years of related experience
- Hands-on production experience in both Azure and AWS.
- Deep expertise in Terraform (modules, workspaces, state, policy as code).
- Strong Kubernetes operational experience (AKS/EKS), including Helm, ingress controllers, ACR/ECR.
- Solid networking fundamentals: VNet/VPC, routing, VPNs, Private Link/Endpoints, ExpressRoute/Direct Connect, load balancers, WAF, DNS.
- Strong identity & access management skills: Entra ID and AWS IAM, SSO/OIDC, secrets management (Key Vault/KMS).
- CI/CD implementation experience with GitHub Actions, Azure DevOps, or Jenkins; security gates and artefact repositories.
- Observability/SRE experience across metrics, logs, tracing, alerting, incident response, and post-mortems.
- Strong scripting abilities (PowerShell, Bash) and OS-level expertise across Linux/Windows.
- Experience with DR patterns (IaC rebuilds), HA architectures, RTO/RPO planning.
Desirable Skills
- M365 Conditional Access (global policies, break-glass, step-up).
- AWS landing zone tooling (Control Tower, IAM Identity Center, account vending/guardrails).
- Ability to read/maintain CloudFormation or Bicep where Terraform is primary.
- Web hosting experience: CDN/WAF (Front Door/CloudFront), TLS/PKI, caching, performance tuning.
- Data fundamentals: S3/Blob lifecycle, RDS/Aurora/SQL MI/Postgres, Redis/ElastiCache/Azure Cache.
- Kubernetes and supply-chain security: admission controls, image signing, SBOM.
Certifications (Preferred)
- Azure: AZ-104, AZ-305, AZ-500 (AZ-700/AZ-400 are a bonus).
- AWS: Solutions Architect – Associate; SA-Pro or DevOps Pro preferred; Security or Advanced Networking is a plus.
- Kubernetes/HashiCorp: CKA, Terraform Associate (CKS is a plus).
- FinOps: FinOps Certified Practitioner (bonus).
Originally posted on Himalayas