发布工程师
Release Engineer
关于 Supabase
Supabase 是基于 Postgres 的开发平台,由开发者为开发者打造。我们提供完整的后端解决方案,包括数据库、认证、存储、边缘函数、实时功能和向量搜索。所有服务深度集成,专为增长而设计。
关于该职位
我们正在寻找一名发布工程师(SRE)加入我们的发布工程团队(属于工程运营部门)——一位生产运维专家,将 SRE 思维带入 Supabase 的发布和运行方式,使部署在大规模下安全、可观测且可恢复。
发布工程的职责已远远超出构建和发布:我们越来越多地负责部署和运行 Supabase 系统的运维可靠性。在这个职位上,你将把我们的部署流水线、预生产信号以及控制平面本身视为生产系统——拥有 SLO、错误预算和轮班值班责任,并且当可靠性出现问题时,其他团队会依赖你。
这不是一个“守门人”角色。你将让可靠路径变得简单:标准化我们的部署方式,对所发布的内容进行监控,并确保当出现问题时,我们可以快速检测并恢复。
你将负责以下内容
在此职位中,你将:
- 根据明确的 SLO 和错误预算,负责 Supabase 部署和发布系统的可靠性,以及它们运行的控制平面
- 将预生产环境转化为可信信号——标准化并监控当前分散、临时的部署流程
- 推动灾难恢复准备,包括从零开始可重复部署环境(理清未记录的密钥、不明确的配置所有权和循环服务依赖)
- 构建并运营关键用户流程的健康状况和 SLO 监控,使用合成测试在客户发现回归问题之前捕捉问题
- 降低与部署相关的事件的平均检测时间和平均恢复时间——这些事件占我们事件负载的很大一部分
- 参与轮班值班,领导无责复盘,并将发现转化为操作手册、警报和自动化以减少重复劳动
- 提高部署的可观测性和可审计性——清晰记录何时、何地、由谁发布了什么
- 记录操作流程——紧急访问路径、访问模型和操作手册——以确保可靠性知识不会成为部落知识
可靠性与运维
- 定义并跟踪 SLA、SLO、错误预算和 DORA 交付指标——在噪音中设置有意义的警报
查看英文原文
ABOUT SUPABASE
Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth.
ABOUT THE ROLE
We're looking for a Release Engineer (SRE) to join our Release Engineering team (part of EngOps) — a production-operations expert who brings an SRE mindset to how Supabase ships and runs, making deploys safe, observable, and recoverable at scale.
Release Engineering's scope has grown well beyond build-and-ship: we increasingly own the operational reliability of the systems that deploy and run Supabase. In this role you'll treat our deployment pipelines, pre-production signal, and the control plane itself as production systems — with SLOs, error budgets, and on-call ownership — and you'll be the person teams lean on when reliability is on the line.
This is not a "gatekeeper" role. You'll make the reliable path the easy path: standardising how we deploy, instrumenting what we ship, and ensuring that when something breaks, we detect it quickly and recover quickly.
WHAT YOU'LL BE RESPONSIBLE FOR
In this role, you'll:
- Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets
- Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows
- Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies)
- Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do
- Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load
- Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil
- Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom
- Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal
RELIABILITY & OPERATIONS
- Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise
- Ensure deployments fail fast and safely when health checks degrade
- Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds
- Partner with product engineering and platform teams to align release practices with reliability and availability targets
YOU MIGHT BE A GOOD FIT IF YOU
- Have 5+ years in SRE, production operations, platform engineering, or release engineering
- Have operated production systems at scale and carried on-call for them
- Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar)
- Have led incident response with tooling like incident.io http://incident.io (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTR
- Operate confidently on AWS (multiple accounts, IAM, VPC) in production
- Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes
- Script and automate to eliminate toil rather than absorb it
- Communicate clearly with both infrastructure specialists and product engineers
- Thrive in async, globally distributed teams
- Are comfortable navigating ambiguity and iterating toward better systems over time
WHAT WE OFFER
- Fully Remote
We hire globally. We believe you can do your best work from anywhere. There are no Supabase offices, but we provide a WeWork membership or co-working allowance you can use anywhere in the world.
- ESOP
Every team member receives ESOP (equity ownership) in the company. We want everyone to share in the upside of what we’re building together.
- Tech Allowance
Use this budget to set up your ideal work environment—laptop, monitor, headphones, or whatever helps you do your best work.
- Health Benefits
Supabase covers 100% of health insurance for employees and 80% for dependents, wherever you are. Your wellbeing and your family’s health are important to us.
- Annual Off-Sites
Once a year, the entire company gathers in a new city for a week of connection, collaboration, and fun. It’s a highlight of our year.
- Flexible Work
We operate asynchronously and trust you to manage your own time. You know what needs to be done and when.
- Professional Development
Every team member receives an annual education allowance to spend on learning—courses, books, conferences, or anything that supports your growth.
ABOUT THE TEAM
Supabase was born-remote and open-source-first. We believe our globally distributed team is our secret weapon in building tools developers love.
- ~400 team members
- 60+ countries
- 20+ languages spoken
- Over $1B raised (including our $500M Series F)
- 540,000+ community members
We move fast, build in public, and use what we ship. If it’s in your project, we probably use it in ours too. We believe deeply in the open-source ecosystem and strive to support—not replace—existing tools and communities.