站点可靠性工程师
Site Reliability Engineer
关于 Supabase
Supabase 是一个基于 Postgres 的开发平台,由开发者为开发者打造。我们提供完整的后端解决方案,包括数据库、认证、存储、边缘函数、实时功能和向量搜索。所有服务深度集成,专为增长而设计。
关于该职位
Supabase 管理数百万个 Postgres 实例,并且正在快速成长。我们在可观测性、发布工程和事件管理方面有强大的团队——并且我们正在将可靠性工作集中到一个专门的 SRE 实践中,将这一学科在平台上统一起来。
你将被嵌入到服务运营团队中,你的主要职责是让每个工程团队更加可靠——不是通过拥有他们的基础设施,而是通过建立实践、框架和反馈循环,让他们自己拥有可靠性。你将在整个组织中工作:有时设定标准,有时与团队配对编程修复问题,有时帮助团队定义他们的错误预算,有时告诉他们错误预算已耗尽。
这个职位适合那些对 SRE 应该如何运作有清晰愿景的人,在异步、快节奏的环境中茁壮成长,影响力比权威更重要。
你将负责
- 与服务团队合作,根据客户体验定义有意义的 SLI 和 SLO,并构建将它们转化为工程决策的错误预算政策
- 负责并持续改进操作就绪评审(ORR)流程——对新服务和重大变更进行评审,涵盖可观测性、警报、操作手册、容量和优雅降级等方面
- 加强从事件到改进的流程:将事后分析结果与操作就绪差距联系起来,识别重复失败模式,并推动系统性修复
- 作为可靠性专家,被团队在架构评审、故障模式分析、依赖关系映射和弹性设计中所引用
- 识别并量化整个组织的运维负担,并构建或倡导自动化以消除它
- 协助团队设计可持续的值班制度:警报质量、升级路径、操作手册覆盖范围和噪音减少
- 跟踪并报告整个组织的操作成熟度,揭示系统性差距并推动整改
你可能适合这个职位如果
- 有 7 年以上 SRE、生产工程或可靠性相关工作的经验,包括塑造 SRE 实践并在工程团队中推动采用的经验
- 有软件工程背景
查看英文原文
ABOUT SUPABASE
Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth.
ABOUT THE ROLE
Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform.
You'll be embedded within Service Operations, and your primary job is to make every engineering team more reliable — not by owning their infrastructure, but by establishing the practices, frameworks, and feedback loops that let them own reliability themselves. You'll work across the org: sometimes setting the standard, sometimes pair-programming a fix, sometimes helping a team define their error budget, sometimes telling them it's exhausted.
This role is ideal for someone who has a strong vision for how SRE should work and thrives in async, fast-paced environments where influence matters more than authority.
WHAT YOU'LL OWN
- Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions
- Own and evolve the Operational Readiness Review (ORR) process — conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation
- Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes
- Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design
- Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it
- Help teams design sustainable on-call practices: alert quality, escalation paths, runbook coverage, and noise reduction
- Track and report on org-wide operational maturity, surfacing systemic gaps and driving remediation
YOU MIGHT BE A GOOD FIT IF YOU
- Have 7+ years of experience in SRE, production engineering, or reliability-focused roles, including experience shaping SRE practices and driving adoption across engineering teams
- Have a software engineering mindset — you write code and build tools, not just configure them
- Have hands-on experience defining and operationalizing SLOs/SLIs at scale, including error budget policies that actually influenced engineering decisions
- Have deep experience with incident response, postmortem facilitation, and turning incident learnings into systemic improvements
- Have worked with large-scale multi-tenant systems (bonus: managed database platforms or Postgres)
- Are proficient with cloud infrastructure (AWS preferred) and infrastructure-as-code (Pulumi preferred, Terraform/CDK also acceptable)
- Communicate clearly and persuasively — this role requires influencing without authority across a distributed org
- Have experience in async or globally distributed teams
- Are energized by making other teams more effective rather than being the one who fixes everything
NICE TO HAVE
- Experience with Kubernetes-based platform operations
- Familiarity with OpenTelemetry, VictoriaMetrics, Grafana, or similar observability tooling
- Experience building developer-facing reliability tooling (SLO dashboards, ORR frameworks, toil tracking, DORA metrics)
WHAT WE OFFER
- Fully Remote
We hire globally. We believe you can do your best work from anywhere. There are no Supabase offices, but we provide a WeWork membership or co-working allowance you can use anywhere in the world.
- ESOP
Every team member receives ESOP (equity ownership) in the company. We want everyone to share in the upside of what we’re building together.
- Tech Allowance
Use this budget to set up your ideal work environment—laptop, monitor, headphones, or whatever helps you do your best work.
- Health Benefits
Supabase covers 100% of health insurance for employees and 80% for dependents, wherever you are. Your wellbeing and your family’s health are important to us.
- Annual Off-Sites
Once a year, the entire company gathers in a new city for a week of connection, collaboration, and fun. It’s a highlight of our year.
- Flexible Work
We operate asynchronously and trust you to manage your own time. You know what needs to be done and when.
- Professional Development
Every team member receives an annual education allowance to spend on learning—courses, books, conferences, or anything that supports your growth.
ABOUT THE TEAM
Supabase was born-remote and open-source-first. We believe our globally distributed team is our secret weapon in building tools developers love.
- ~400 team members
- 60+ countries
- 20+ languages spoken
- Over $1B raised (including our $500M Series F)
- 540,000+ community members
We move fast, build in public, and use what we ship. If it’s in your project, we probably use it in ours too. We believe deeply in the open-source ecosystem and strive to support—not replace—existing tools and communities.