高级基础设施工程师,SRE
Senior Infrastructure Engineer, SRE
**关于 Rocket Money 🔮**
Rocket Money 的使命是赋能人们过上最好的财务生活。Rocket Money 为会员提供对其财务状况的独特理解,以及一系列有价值的的服务,帮助他们节省时间和金钱——最终让他们在财务旅程中获得优势。
**关于团队 🤹**
我们正在寻找一位高级基础设施工程师、SRE 来带领我们平台的可靠性和运营发展。我们在生产环境中运行数百个服务,这些服务使我们能够处理数十亿次交易,消耗多个 TB 的数据,并每天生成数亿条日志,而我们的可靠性实践需要随着我们不断增长的规模而演进。这包括:
- 构建和改进系统的可靠性和弹性
- 为最关键的服务和用户流程建立 SLI、SLO 和错误预算,并定期与相关团队进行审查
- 负责并不断发展我们的灾难恢复策略:恢复目标、故障转移和恢复路径,以及定期演练以证明其有效性
- 与产品工程团队合作,使他们能够自主拥有和运营自己的服务,并具备反映真实用户体验的指标
- 演进我们的可观测性平台和标准,涵盖指标、追踪和日志:包括仪器化路径、警报质量及可观测性成本
- 加强我们的事件管理实践:调整通知阈值,保持操作手册最新,并跟进事后分析中的行动项
- 在日常云基础设施工作中贡献力量,同时专注于可靠性专业领域——基础设施部署、平台待办事项,以及共享的值班轮班(每六周值班一周)
你将加入云基础设施团队,并与工程和内部支持团队合作推动这项工作。
我们帮助数百万人改善他们的财务生活,而这个职位确保我们能够继续可靠且可扩展地做到这一点。
**关于你 🦄**
- 你有 5 年以上的云或基础设施工程经验,其中大量时间用于大规模的可靠性和生产运维
- 你定义过真实生产服务的 SLI 和 SLO,并能谈论由此带来的变化。哪些问题得到了修复,哪些被降级,以及第一次你犯了哪些错误
- 你在生产环境中使用过可观测性平台;强烈推荐使用 Datadog
查看英文原文
**ABOUT ROCKET MONEY 🔮**
Rocket Money’s mission is to empower people to live their best financial lives. Rocket Money offers members a unique understanding of their finances and a suite of valuable services that save them time and money – ultimately giving them a leg up on their financial journey.
**ABOUT THE TEAM 🤹**
We're looking to expand our Cloud Infrastructure team with a Senior Infrastructure Engineer, SRE to lead the reliability and operational evolution of our platform. We run hundreds of services in production, which enable us to process billions of transactions, consume multiple terabytes of data, and produce hundreds of millions of logs per day, and our reliability practice needs to evolve to match our growing scale. This includes:
- Building and improving the reliability and resiliency of our systems and services
- Establishing SLIs, SLOs, and error budgets for our most critical services and user journeys, and reviewing them regularly with the teams that own them
- Owning and evolving our disaster recovery strategy: recovery objectives, failover and restore paths, and regular exercises that prove they work
- Partnering with product engineering teams so they can own and operate their own services, with metrics that reflect real user experience
- Evolving our observability platform and standards across metrics, tracing, and logs: including instrumentation paved roads, alert quality, and observability cost
- Strengthening our incident practice: tuning paging thresholds, keeping runbooks current, and following through on postmortem action items
- Contributing to day-to-day Cloud Infrastructure work alongside your reliability specialty — infrastructure build-outs, platform backlog, and a shared on-call rotation (1 week out of every 6 weeks)
You'll join the Cloud Infrastructure team and partner with engineering and internal support teams to drive this work.
We support millions of people to improve their financial lives, and this role ensures we can continue to do so reliably and at scale.
**ABOUT YOU 🦄**
- You have 5+ years of hands-on cloud or infrastructure engineering experience, with substantial time spent on reliability and production operations at scale
- You have defined SLIs and SLOs for real production services, and can talk about what changed as a result. What got fixed, what got deprioritized, and what you got wrong the first time
- You have hands-on experience with an observability platform in production; Datadog strongly preferred
- You're comfortable writing code (Python, Go, TypeScript, or similar) for internal tooling, production debugging, and automation
- You write production Terraform and are comfortable in AWS, and when production breaks you can find the problem and fix it
- You have built or operated a disaster recovery plan: you set the recovery goals, wrote the failover and restore steps, and ran the drills that proved it works
- You have been on-call for services you helped build, and you have opinions about what makes an alert worth waking someone for
- You prefer giving teams paved roads and good defaults over mandates, so they can own their own instrumentation
##### **BONUS POINTS**
- You have led a reliability or observability modernization project where you defined the vision, approach, and delivered the implementation
- You have built internal tooling, libraries, or instrumentation standards that made it easier for other teams to operate their services well
- You have run game days, chaos experiments, or DR exercises, and fixed the problems they uncovered
- You have cut observability spend while keeping the coverage you needed
**WE OFFER 💫**
- Health, Dental & Vision Plans
- Competitive Pay
- 401k Matching
- Unlimited PTO
- Lunch daily _(in-office only)_
- Snacks & Coffee _(in-office only)_
- Commuter benefits _(in-office only)_
Additional information: Salary range of $150,000 - $185,000/year + bonus + benefits. Base pay offered may vary depending on job-related knowledge, skills, and experience.
_This job description is an outline of the primary responsibilities of this position and may be modified at the discretion of the company at any time. Decisions related to employment are not based on race, color, religion, national origin, sex, physical or mental disability, sexual orientation, gender identity or expression, age, military or veteran status or any other characteristic protected by state or federal law. The company provides reasonable accommodations to qualified individuals with disabilities in accordance with applicable state and federal laws._
_The information regarding compensation and other benefits included in this paragraph is the company’s current, good faith estimate at the time of posting. [Compensation and benefits are subject to modification from time to time as the Company, in its sole and exclusive discretion, deems appropriate.] The Company may determine during its future reviews of the proposed compensation and benefits provided for this position, that the compensation and benefits for such position should be reduced. In no event will the Company reduce the compensation for the position to a level below the applicable jurisdictional minimum wage rate for the position. Los Angeles County and San Francisco Candidates only: qualified applicants with arrest or conviction records will be considered for employment per the Fair Chance Ordinance and the Fair Chance Initiative for Hiring._