高级软件工程师,DevOps
Senior Software Engineer, DevOps
关于Atria:
Atria健康研究所是一家以预防和长寿为重点的会员制初级和专科医疗服务机构。我们汇聚了一支跨学科的知名医生团队,为Atria会员及其家庭提供主动、预防和精准的医疗护理。所有服务,包括初级护理、高级筛查和诊断、紧急护理、专科护理、全天候上门服务和影像检查,均包含在会员年度费用中。
我们的使命是通过实时将科学转化为医学,使每个人的健康寿命和寿命相等,同时将人性带回医疗保健。提供如此强大、个性化和预防性的医疗保健是一项复杂的任务,需要全团队对卓越的承诺。
在2022年成功开设纽约旗舰研究所,并于2024年扩展到南佛罗里达后,我们现在将于2026年春季在西海岸推出洛杉矶研究所。
Atria Health正在寻找一名高级软件工程师加入我们的DevOps团队,帮助设计、构建和运营基础设施、部署流水线和可观测性系统,这些系统是工程团队每天依赖的基础。
这是一个个人贡献者角色,重点包括:
- 从设计到部署、采用和迭代,全程负责基础设施、CI/CD和自动化项目
- 设定模式和标准,确保我们在扩展过程中保持系统的可靠性、可观测性和安全性
- 减少运维负担,提升整个工程团队的开发效率
你将与技术负责人和其他DevOps工程师合作,塑造我们的基础设施即代码、构建和发布流程,以及监控和安全态势——并帮助定义其他人所依赖的标准。你将与平台、数据工程和领域产品团队(临床体验、会员体验和护理交付)紧密合作,预判他们的需求,使他们的生产路径更加顺畅。你将 deeply 关注我们如何构建以及实现什么目标,同时加速客户团队的工作。
基础设施与自动化
- 在Google Cloud Platform上使用Terraform设计、构建和维护云基础设施,并帮助定义团队所依赖的模式和标准。
- 负责并改进GitHub Actions中的CI/CD流水线,使部署快速、安全且可重复。
- 识别并消除运维负担的来源,建立自动化解决方案
查看英文原文
About Atria:
The Atria Health Institute is a membership-based primary and specialty health care practice with a focus on prevention and longevity. We bring together a multidisciplinary team of renowned physicians to provide proactive, preventive, and precision-based care for Atria members and their families. All care, including primary care, advanced screening and diagnostics, urgent care, specialty care, 24/7 home visits, and imaging is included in members’ annual fee.
Our mission is to make healthspan and lifespan equal for all by translating science into medicine in real-time, all while bringing humanity back into health care. Delivering such robust, personalized, and preventive health care is complex and requires a team-wide dedication to excellence.
After successfully opening our flagship Institute in New York in 2022 and expanding to South Florida in 2024, we are now bringing the Atria experience to the West Coast with the launch of our Los Angeles Institute in late spring 2026.
Atria Health is seeking a Senior Software Engineer for our DevOps team to help design, build, and operate the infrastructure, deployment pipelines, and observability that the rest of engineering relies on every day.
This is an individual contributor role focused on:
- Owning infrastructure, CI/CD, and automation initiatives end to end; from design through rollout, adoption, and iteration
- Setting the patterns and standards that keep our systems reliable, observable, and secure as we scale
- Reducing operational toil and raising the bar on developer productivity across engineering
You'll partner with the Tech Lead and other DevOps engineers to shape our infrastructure-as-code, build and release process, and monitoring and security posture — and you'll help define the standards others build on. You'll work closely with the Platform, Data Engineering, and domain product teams (Clinical Experience, Member Experience, and Care Delivery) to anticipate their needs and make their path to production smoother. You'll care deeply about both how we build and what we achieve, while working to accelerate the work of our client teams.
Infrastructure & Automation
- Design, build, and maintain cloud infrastructure on Google Cloud Platform using Terraform, and help define the patterns and standards the team builds on.
- Own and improve CI/CD pipelines in GitHub Actions to make deployments fast, safe, and repeatable.
- Identify and eliminate sources of operational toil, building automation and tooling that scales across engineering.
Reliability & Observability
- Establish meaningful monitoring, dashboards, and alerting in Datadog, and drive down alert noise across the team's systems.
- Lead incident response within the on-call rotation, and drive postmortems and follow-ups that improve uptime for our applications.
- Define and track Service Level Objectives (SLOs) for our core infrastructure and build systems, and hold the team to them.
Developer Enablement
- Partner with product engineering teams on deployment pipelines, environment issues, and build troubleshooting, and proactively remove recurring friction.
- Own preview and staging environments, including reliable data sync, masking, and cleanup routines so teams can test against realistic data.
- Improve developer experience through better tooling, clear runbooks, and documentation.
Quality & Collaboration
- Write well-tested, well-reviewed infrastructure code, and lead design reviews and RFCs with a pragmatic, operations-focused perspective.
- Design systems and changes to meet the team's goals, weighing tradeoffs across reliability, performance, and security.
- Give thoughtful code reviews and mentor other engineers through pairing, knowledge sharing, and documentation.
Tech Stack
- Languages: TypeScript, Python, Bash
- Infrastructure: Google Cloud Platform, Terraform, Cloudflare
- CI/CD: GitHub Actions
- Databases: MySQL, PostgreSQL, Redis
- Observability & Incident Management: Datadog, Sentry, Rootly
- Integrations: Athena EMR, wearable platforms, third-party healthcare APIs
Requirements
Core Experience
- ~5+ years of professional experience in DevOps, SRE, infrastructure, or backend engineering in production environments.
- Hands-on experience designing and operating infrastructure in at least one cloud provider (ideally Google Cloud Platform).
- Track record of owning and shipping automation, pipelines, or infrastructure that made a team measurably more productive or reliable.
- An enthusiasm for developer productivity and making our teams as impactful as possible.
Technical Skills
- Deep experience with infrastructure-as-code (ideally Terraform) and building CI/CD pipelines (ideally GitHub Actions).
- Proficient software engineering ability, and strong command of Linux.
- Strong instincts for monitoring and observability, and confident debugging across logs, traces, and metrics.
- Solid experience with relational databases (MySQL, PostgreSQL) and containerized workloads.
- Strong grounding in reliability, performance, and security fundamentals, with the judgment to make sound tradeoffs.
Nice to Have
- Experience in healthcare, digital health, or other regulated domains (HIPAA, PHI, SOC 2, etc.).
- Experience with containers and orchestration (Docker, Kubernetes).
- Exposure to leading incident response, on-call, and postmortem practices.
- Experience with database migrations or managing multiple environments at scale.
Originally posted on Himalayas