远程工作雷达

资深站点可靠性工程师

Principal Site Reliability Engineer

开发工程职能支持限定地区(需当地身份)
公司Tandem Diabetes Care
薪资$165,000 - $185,000/年
工作地点United States
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

与我们共同成长:
Tandem Diabetes Care 通过一种积极不同的体验,为糖尿病患者、他们的亲人以及医疗保健提供者创造新的可能。我们希望您与我们携手“每天创新”,以“以人为本”为原则,并采用“不走捷径”的方法,使我们成为糖尿病技术行业的领导者。
保持卓越:
Tandem Diabetes Care 以制造和销售 Tandem Mobi 系统和带有 Control-IQ+ 技术的 t:slim X2 胰岛素泵而自豪——这是一种先进的预测算法,可自动输送胰岛素。但我们远不止于此。我们以用户为中心的设计、开发和支持方法,为使用胰岛素的人提供创新的产品和服务。由于我们许多团队成员自己患有糖尿病,或有亲人受到糖尿病影响,这项工作对我们来说意义重大,我们致力于这一事业。了解更多请访问 tandemdiabetes.com
日常职责:
高级站点可靠性工程师(SRE)负责公司生产系统的可靠性、可用性和性能。该职位负责日常生产支持和事件响应,并逐步用工程化的 SRE 实践取代被动应对:SLO、可观测性、值班设计、操作手册和自动化。同时,该职位与跨职能团队合作,推进基础设施自动化(Terraform/IaC)、CI/CD 可靠性、灾难恢复以及安全和合规准备。该职位与一个包括海外咨询合作伙伴的分布式团队合作,期望提升他们的能力和独立性;成功不仅取决于高级 SRE 个人的贡献,也取决于团队在没有高级 SRE 的情况下能做什么。Tandem 的高级站点可靠性工程师(SRE)还负责:
生产支持与事件管理

  • 领导日常生产支持:接收、分类、优先级排序、升级、队列健康状况和变更执行。
  • 在分布式团队中建立一致的支持实践,包括班次交接、工单质量标准和明确的问题所有权。
  • 全程领导事件管理:事件指挥、利益相关者沟通以及无责复盘并跟踪纠正措施直至关闭。
  • 参与并协调生产安全事件的响应活动,与安全团队合作遏制威胁并恢复系统。
查看英文原文

GROW WITH US:
Tandem Diabetes Care creates new possibilities for people living with diabetes, their loved ones, and their healthcare providers through a positively different experience. We’d love for you to team up with us to “innovate every day,” put “people first,” and take the “no-shortcuts” approach that has propelled us to become a leader in the diabetes technology industry.
STAY AWESOME:
Tandem Diabetes Care is proud to manufacture and sell the Tandem Mobi system and t:slim X2 insulin pump with Control-IQ+ technology — an advanced predictive algorithm that automates insulin delivery. But we’re so much more than that. Our company’s human-centered approach to design, development, and support delivers innovative products and services for people who use insulin. Because many of our own team members live with diabetes, or have a loved one impacted by diabetes, the work is personal, and we are committed to the cause. Learn more at tandemdiabetes.com
A DAY IN THE LIFE:
The Principal Site Reliability Engineer (SRE) is responsible for the reliability, availability, and performance of the company's production systems. This role leads day-to-day production support and incident response and progressively replaces reactive firefighting with engineered SRE practice: SLOs, observability, on-call design, runbooks, and automation. It also advances infrastructure automation (Terraform/IaC), CI/CD reliability, disaster recovery, and security and compliance readiness in partnership with cross-functional teams. The role works with a distributed team that includes offshore consulting partners and is expected to raise their capability and independence; success is measured as much by what the team can do without the Principal SRE as by what the Principal SRE delivers personally.The Principal Site Reliability Engineer (SRE)'s at Tandem are also responsible for:
Production Support & Incident Management

  • Leads day-to-day production support: intake, triage, prioritization, escalation, queue health, and change execution.
  • Establishes consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and clear ownership of open issues.
  • Leads incident management end-to-end: incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure.
  • Participates in and coordinates response activities for production security incidents, partnering with Security teams to contain threats, restore services, validate controls, and implement corrective actions.
  • Owns on-call strategy: rotation design, escalation paths, alert tuning, and tooling (e.g., PagerDuty, New Relic), with explicit goals of reducing alert fatigue and building coverage that works across time zones. Participates as a senior escalation tier for high-severity incidents.
  • Builds runbooks that standardize response to common failure modes and enable first-line resolution by engineers who did not build the system.

Reliability Engineering & Observability

  • Defines and owns SLIs and SLOs for critical services, using them to guide monitoring, alerting, and reliability priorities.
  • Drives systemic reduction of MTTD and MTTR through better instrumentation, alerting, diagnostics, and automation.
  • Converts recurring support burden into permanent fixes, automation, or documentation rather than absorbing it as ongoing manual work.
  • Establishes and maintains technology currency and lifecycle management practices for production platforms, ensuring cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components remain supported, secure, and aligned with organizational standards. Proactively identifies and mitigates End-of-Life (EOL), End-of-Support (EOS), and technology obsolescence risks.
  • Owns business continuity and disaster recovery readiness for production platforms, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and adherence to defined RTO/RPO objectives.

Infrastructure Automation & CI/CD

  • Leads infrastructure automation with Terraform, ensuring infrastructure is version-controlled, modular, reusable, and auditable, working within and improving existing patterns where they are sound.
  • Eliminates toil through automation, reducing manual, repetitive operational work across the team.
  • Adds reliability guardrails to CI/CD pipelines (automated rollback, change-risk checks, progressive delivery) with the teams that own them.

Operational Governance and Compliance

  • Maintains production systems in accordance with applicable regulatory and compliance requirements, ensuring audit readiness and the ongoing effectiveness of access management, change management, logging, vulnerability remediation, patch management, and other operational controls.
  • Maintains documentation and evidence supporting business continuity and disaster recovery controls, including recovery testing results, remediation plans, and audit artifacts.
  • Partners with Security, Quality, and Compliance teams to maintain regulatory compliance, support internal and external audits, and drive timely resolution of audit findings and remediation activities.

Leadership, Mentorship & Collaboration

  • Actively grows the capability of SRE and DevOps engineers, including consulting-partner engineers, through pairing, design and code review, and incident debriefs.
  • Creates conditions in which junior and contract engineers raise concerns and disagree openly, recognizing that silent agreement is a production risk.
  • Communicates in a documentation-first manner suited to a distributed, multi-time-zone team, so context is durable rather than held by one person.
  • Partners with software engineering, QA, and architecture to embed reliability and operability into the development lifecycle rather than only after incidents.

Contributing Responsibilities (partners with other owners):

  • Informs capacity planning and scaling strategy with the Test team (who own load testing and modeling) and introduces proactive resilience testing such as game days once baseline support and observability are stable.
  • Partners with application, architecture, and business stakeholders to align business continuity and disaster recovery capabilities with application requirements and recovery objectives.
  • Supports cloud cost optimization: rightsizing, reserved capacity, and observability spend governance.
  • Ensures work complies with company policies, including Privacy/HIPAA and other regulatory, legal, and safety requirements. Other duties as assigned.

WHEN & WHERE YOU’LL WORK:
Remote: This position is fully remote and open to candidates within the United States. Equipment for the role will be provided and training will occur virtually.
WHAT YOU’LL NEED:

  • We are looking for a forward-thinking professional who actively leverages AI to enhance, automate, and optimize internal tooling and system integrations. In this role, you won't just use existing systems—you will use AI to actively build smarter workflows and bridge gaps between our platforms.
  • Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events.
  • Strong grounding in SRE principles: SLIs/SLOs, blameless postmortems, toil reduction, and treating reliability as an engineering discipline rather than a purely operational function.
  • Demonstrated experience owning on-call strategy, including rotation design, alert tuning, and escalation.
  • Expertise with Terraform or comparable IaC at scale: module design, state management, and policy-as-code guardrails.
  • Hands-on experience building CI/CD pipelines with reliability guardrails using tools such as GitHub Actions, Octopus Deploy, or Azure DevOps.
  • Deep experience with at least one major cloud platform (AWS, Azure, or GCP; [ preferred]) and with containerization and orchestration (Docker/Kubernetes).
  • Working knowledge of observability tooling (e.g., Prometheus, Grafana, Datadog, CloudWatch, ELK/OpenSearch) and using it to drive SLO-based alerting.
  • Experience designing and testing disaster recovery, including backup/restore, failover, and RTO/RPO validation.
  • Working knowledge of cloud security and compliance practices (IAM, network segmentation, encryption, vulnerability and patch management) and cloud cost optimization.
  • Proficiency in at least one scripting or programming language (e.g., Python, Go, Bash) for automation and tooling.
  • Experience in FDA and ISO regulated industries and with agile methodologies preferred.

EXTRA AWESOME:

  • B.S. in Computer Science or an equivalent combination of education and applicable job experience, including technical school training and certifications in networks, servers, and cloud infrastructure; demonstrated production experience weighs more heavily than degree.
  • Relevant cloud certifications (e.g., AWS/Azure/GCP Professional or Architect level) preferred.
  • 10+ years in Site Reliability Engineering, DevOps, or infrastructure engineering, including:
  • Leading production support, incident management, and postmortem practice.
  • Managing production cloud infrastructure as code (Terraform or equivalent) at scale.
  • Operating and improving CI/CD pipelines in a production environment.
  • Designing, executing, and testing disaster recovery plans.
  • Partnering with security teams on production security and compliance posture.
  • 2+ years mentoring or technically leading other engineers, including engineers who are remote, offshore, or contracted through a partner organization.

COMPENSATION & BENEFITS:
The starting base pay range for this position is $165,000 to $185,000 annually. Base pay will vary based on job-related knowledge, skills, experience and may also fluctuate depending on candidate’s location and the overall job market. In addition to base pay, Tandem offers a competitive compensation package that includes bonus and a robust benefits package.
Tandem offers health care benefits such as medical, dental, vision available your first day, as well as health savings accounts and flexible saving accounts. You’ll also receive 11 paid holidays per year, a minimum of 20 days of paid time off (with accrual starting on day 1) and you will have access to a 401k plan with company match as well as an Employee Stock Purchase plan. Learn more about Tandem’s benefits here!
YOU SHOULD KNOW:
Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable state and local Fair Chance laws and regulations. A conditional offer of employment from Tandem is contingent upon successful completion of a pre-employment screening process comprised of a drug test (excluding marijuana) and background check, which includes a review of criminal history information.
Tandem has good cause to conduct a review of criminal history information of candidates for this position, as this role may involve access to proprietary, sensitive and/or confidential information, including customer protected health information. This review is required to ensure that individuals in such roles uphold high standards of trust and integrity so as to protect the interests of our customers, employees, and stakeholders.
WHY YOU’LL LOVE WORKING HERE:
At Tandem, we believe joy fuels excellence. That's why we've built a workplace that celebrates your achievements and supports your well-being. Our team thrives on pushing boundaries and fostering growth, all while maintaining a spirit of fun and camaraderie. This is just one of the ways we stay awesome! Explore the benefits and reasons to love Tandem at .
BE YOU, WITH US!
We embrace the value that every single one of us brings to the table. But sometimes we forget that when we don’t meet 100% of a job description’s criteria – maybe you’re feeling that way right now? We encourage you to apply anyway. Because we want you to be you, with us.
Tandem is firmly committed to being an equal opportunity employer and does not discriminate on the basis of age, disability, sex, race, religion or belief, gender identity or expression, marriage/civil partnership, pregnancy/maternity, or sexual orientation. We are an inclusive organization, and we welcome applications from a wide range of candidates. Selection for roles will be based on individual merit alone.
REFERRALS:
We love a good referral! If you know someone who would be a great fit for this position, please share!
APPLICATION DEADLINE:
The position will be posted until a final candidate is selected for the requisition or the requisition has a sufficient number of applications.
Make a move that matters. Join Tandem Diabetes Care, where we're turning challenges into triumphs every day and where your talents will help shape a healthier, happier tomorrow.
Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

临床糖尿病销售专员

Tandem Diabetes CareNetherlandsFull Time今天
市场运营限定地区(需当地身份)日间重叠约 2 小时,需偶尔早起或晚睡

高级营销技术策略师

Tandem Diabetes CareUnited States$89,900 - $112,250/年Full Time2 天前
开发工程市场运营职能支持限定地区(需当地身份)

← 返回全部职位