资深站点可靠性工程师(SRE)
Principal Site Reliability Engineer (SRE)
Symmetrio正在为一家快速发展的医疗技术公司招聘首席站点可靠性工程师(SRE),该公司专注于先进的医疗技术解决方案。
该职位将发挥关键作用,确保支持美国各地医疗提供者的任务关键型SaaS平台的可靠性、可扩展性、安全性和性能。理想的候选人应具备云基础设施专业知识、应用故障排除经验、生产运维领导力以及面向客户的解决技术问题的能力。
理想的候选人应同样擅长调查应用级问题、排查AWS网络和基础设施、领导生产事件响应工作,并与开发团队合作提升运营卓越性。
这是一项全远程职位,薪资范围为18万美元至20万美元,具体根据经验而定。
职责
- 作为美国客户环境生产可靠性的主要技术负责人。
- 调查并解决涉及Web应用、API、后端服务、数据管道、云基础设施和客户集成的复杂问题。
- 领导生产事件响应工作,协调跨职能团队恢复服务并最小化客户影响。
- 执行根本原因分析并推动纠正措施,以提高系统的长期稳定性和弹性。
- 与软件工程和平台团队合作,识别重复出现的可靠性风险并实施可持续解决方案。
- 设计、配置和验证安全的客户连接解决方案,包括站点到站点VPN、传输网关集成、路由配置和安全网络路径。
- 通过排查连接问题并确保一致的实施流程,支持客户上线计划。
- 通过改进监控、日志、警报、追踪和操作仪表板来增强平台可观测性。
- 参与CI/CD、基础设施自动化和部署流程,以提高发布安全性和操作一致性。
- 开发支持事件响应、故障排除、上线和系统监控活动的操作工具。
- 与工程领导层合作,提升云架构、可扩展性、安全性和操作准备度。
- 与面向客户的团队合作,沟通技术方案和解决方案。
查看英文原文
Symmetrio is recruiting a Principal Site Reliability Engineer (SRE) for our customer, a rapidly growing healthcare technology organization focused on advanced healthcare technology solutions.
This individual will play a critical role in ensuring the reliability, scalability, security, and performance of a mission-critical SaaS platform supporting healthcare providers across the United States. The ideal candidate will possess a unique blend of cloud infrastructure expertise, application troubleshooting experience, production operations leadership, and customer-facing technical problem-solving skills.
The ideal candidate will be equally comfortable investigating application-level issues, troubleshooting AWS networking and infrastructure, leading production incident response efforts, and collaborating with development teams to improve operational excellence.
This is a fully remote position with a salary range of $180,000–$200,000, depending on experience.
Responsibilities
- Serve as the primary technical owner for production reliability across U.S. customer environments.
- Investigate and resolve complex issues spanning web applications, APIs, backend services, data pipelines, cloud infrastructure, and customer integrations.
- Lead production incident response efforts, coordinating cross-functional teams to restore service and minimize customer impact.
- Perform root cause analysis and drive corrective actions that improve long-term system stability and resilience.
- Partner with software engineering and platform teams to identify recurring reliability risks and implement sustainable solutions.
- Design, configure, and validate secure customer connectivity solutions including Site-to-Site VPNs, Transit Gateway integrations, routing configurations, and secure network paths.
- Support customer onboarding initiatives by troubleshooting connectivity challenges and ensuring consistent implementation processes.
- Enhance platform observability through improvements in monitoring, logging, alerting, tracing, and operational dashboards.
- Contribute to CI/CD, infrastructure automation, and deployment processes that improve release safety and operational consistency.
- Develop operational tooling that supports incident response, troubleshooting, onboarding, and system monitoring activities.
- Collaborate with engineering leadership to improve cloud architecture, scalability, security, and operational readiness.
- Partner with customer-facing teams to communicate technical issues, remediation plans, and reliability improvements in a clear and effective manner.
- Support compliance, security, and risk management initiatives within highly regulated healthcare environments.
Requirements
- 6+ years of hands-on experience supporting and managing AWS-based production environments.
- 4+ years of experience supporting web applications and backend services (Python/Django experience strongly preferred).
- Experience with AWS networking technologies including VPCs, Site-to-Site VPNs, Transit Gateways, routing, NAT gateways, and security groups.
- Strong experience with Terraform and infrastructure-as-code deployment practices.
- Experience with containerized environments including ECS, Fargate, Kubernetes, or similar technologies.
- Experience building and supporting CI/CD pipelines and release automation processes.
- Familiarity with monitoring and observability platforms such as Datadog, CloudWatch, Sentry, Grafana, or similar tools.
- Experience leading production incidents, outage management, and root cause analysis initiatives.
- Exposure to Windows Server environments, Active Directory, Kerberos, and enterprise infrastructure concepts is preferred.
- Healthcare technology, healthcare SaaS, clinical software, or other regulated industry experience is highly preferred.
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field preferred.
Benefits
- Health Care Plan (Medical, Dental & Vision)
- Retirement Plan (401k, IRA)
- Paid Time Off (Vacation, Sick & Public Holidays)
Originally posted on Himalayas