平台运维工程师
Platform Operations Engineer
_此职位为美国境内的全远程岗位。_
Turquoise Health 正从零开始构建平台工程团队,这是你参与塑造它的机会。作为我们的首位平台运维工程师,你将帮助建立支撑真实医疗成果的基础设施的构建、部署和运营基础。
这是一个在新团队中具有高度自主权的角色。你将帮助建立工程团队所依赖的实践、工具和标准,从可观测性和可靠性到部署流程和可扩展性。如果你热衷于解决全新的问题,能够应对不确定性,并且有动力在被要求之前就进行自动化和简化,我们希望与你一起打造伟大的成果。
### 职责:
- **支持平台基础设施** - 协助管理并扩展我们在 Amazon EKS 上的容器环境,使用 ArgoCD 实现 GitOps 流程,并通过 GitHub Actions 维护 CI/CD 流水线,确保部署快速、一致且自动化
- **构建可靠性** - 定义并跟踪 SLI 和 SLO,领导事件响应包括值班轮班、根本原因分析和事后总结,并参与灾难恢复计划以保持系统高可用性
- **推动可观测性** - 使用 Datadog、Sentry 和 CloudWatch 设计和维护监控和日志堆栈,使工程团队能够在问题影响用户前清晰地看到系统健康状况和性能
- **塑造平台的未来** - 参与架构决策,构建内部工具和自助服务流程,使平台更易于操作,并对如何扩展和演进我们的基础设施做出有意义的贡献
### **你将带来的:**
- 3年以上 SRE、DevOps 或云基础设施相关经验
- 熟练使用核心 AWS 服务(VPC、IAM、EKS、RDS),并具备扎实的云网络和安全最佳实践知识
- 精通使用 Infrastructure as Code 工具如 Terraform、CloudFormation 或 Crossplane
- 熟练使用 GitHub 和 GitHub Actions 作为 CI/CD 和自动化流水线的核心组件——而不仅仅是用于源代码控制
- 具有在生产环境中运行 Kubernetes 集群的经验,并能通过 GitOps 流程(ArgoCD/Flux)和 Helm Charts 管理应用部署
- 熟练使用可观测性工具如 Datadog、Sentry、CloudWatch、Grafana,包括构建告警、仪表板和日志流水线
- 相关经验
查看英文原文
_This is a fully remote role within the United States._
Turquoise Health is building its Platform Engineering function from the ground up and this is your chance to shape it. As our first Platform Operations Engineer, you'll help lay the foundation for how we build, deploy, and operate infrastructure that supports real-world healthcare outcomes.
This is a high-ownership role on a new team. You'll help establish the practices, tooling, and standards that the rest of engineering will rely on, from observability and reliability to deployment workflows and scalability. If you're energized by greenfield problems, comfortable navigating ambiguity, and motivated to automate and simplify before being asked, we'd love to build something great together.
### Responsibilities:
- **Support the Platform Infrastructure** - Help manage and scale our container environment on Amazon EKS, implement GitOps workflows using ArgoCD, and maintain CI/CD pipelines through GitHub Actions to ensure that deployments are fast, consistent, and automated
- **Build for Reliability**- Define and track SLIs and SLOs, lead incident response including on-call rotations, root cause analysis, and post-mortems, and contribute to disaster recovery planning to keep our systems highly available
- **Drive Observability**- Design and maintain our monitoring and logging stack using Datadog, Sentry, and CloudWatch — giving engineering teams clear visibility into system health and performance before problems reach users
- **Shape the Platform's Future** - Collaborate on architectural decisions, build internal tooling and self-service workflows that make the platform easier to operate, and contribute meaningfully to how we scale and evolve our infrastructure
### **What You’ll Bring:**
- 3+ years in SRE, DevOps, or Cloud Infrastructure
- Confident working with core AWS services (VPC, IAM, EKS, RDS) and a strong understanding of cloud networking and security best practices
- Expert in using Infrastructure as code with Terraform, CloudFormation, or Crossplane
- Proficient with GitHub and GitHub Actions as a core component of your CI/CD and automation pipelines- not just for source control
- Experienced with running Kubernetes clusters in production and managing application deployments through GitOps workflows (ArgoCD/Flux) and Helm Charts
- Proficient with observability tooling such as Datadog, Sentry, CloudWatch, Grafana to include building alerts, dashboards, and log pipelines
- Experience writing solid Python scripts to glue systems together, automate infrastructure tasks, or handle custom workflows
- Comfortable working independently in a remote setup, asking questions when needed, and keeping momentum without being micromanaged
- Bachelor’s degree in Computer Science, Engineering, or equivalent experience. _Happy to work with strong candidates with non-traditional educational backgrounds_
### **Nice to haves:**
- Certifications: AWS, Kubernetes, Terraform or Python
### **Benefits:**
- Competitive pay with equity options
- Stellar health care plan options (Medical, Dental & Vision), with FSA, DCFSA, & HSA options
- Company-sponsored disability & life insurance
- Unlimited PTO
- 401(k) + 4% Matching
- Fully remote work + flexible working hours
- $750 work-from-home setup budget
- Paid biannual in-person company summits
- Quarterly $150 co-hanging stipend to meet up with coworkers
- Monthly $100 health and wellness benefit
- Generous paid family leave
- Annual $1,200 learning & development stipend
## **About Turquoise Health**
Turquoise Health is a Series C price transparency platform for finance leaders across healthcare. Backed by a16z, Oak HC/FT, Adams Street, Yosemite, Bessemer Venture Partners, and others, we power price transparency for 300+ enterprise organizations and are building the infrastructure for a more open, efficient healthcare marketplace. We're a remote-first, US-based team that values transparency, empathy, inclusivity, creativity, and ownership.
We operate on US business hours and work with clients entirely based in the US. For this role, we are seeking US-based candidates.
_We strongly encourage BIPOC, people with disabilities, and LGBTQIA+ folks to apply for any open roles of interest. Healthcare affects all people differently, but it significantly affects those in underserved communities. With a robust, diverse team, we are stronger and better equipped to change the future of healthcare for all._
### **Disability Accommodation Email**
Turquoise is committed to providing reasonable accommodations to applicants and employees with disabilities. Please tell us if you require a reasonable accommodation to apply for a job or to perform your job. Examples of reasonable accommodation include making a change to the application process or work procedures, providing documents in an alternate format, using a sign language interpreter, or using specialized equipment. If you require assistance or an accommodation with the hiring process, please contact recruiting@turquoise.health