数据库可靠性工程师
Database Reliability Engineer
职位描述
我们每周有数千万的访问量,致力于为用户提供行业最佳体验。你将负责管理我们现有和新数据库基础设施,支持多个平台,并推动向新地区的发布。四年前我们从零开始构建系统,因此你将使用最新技术,无需担心老旧技术。同时,你也有自由去研究并根据你和团队的判断将新技术引入我们的技术栈。
我们的技术栈
- 关系型数据库:AWS Aurora MySQL/PostgreSQL
- 非关系型数据库:MongoDB Atlas, AWS DocumentDB
- 缓存层:AWS ElastiCache, Valkey
- 消息队列:Apache RocketMQ, RabbitMQ
- 搜索引擎:ElasticSearch, Mongo Atlas
- OLAP:Redshift
- 监控与告警:Prometheus, Grafana, Loki, Alert Manager
- 备份与恢复:AWS Backup, 点时间恢复, 跨区域复制
- 数据库代理:RDS Proxy
- 基础设施即代码:Terraform, Ansible
- CI/CD 集成:Jenkins, ArgoCD, Github Action, Helm
- 网络与安全:AWS VPC, 安全组, AWS Secrets Manager
- 编程语言:Golang, Python
- 容器化:Kubernetes
你将负责的工作
- 与 DBRE 和 DevOps 专业人员组成团队工作
- 设计、实施和维护高可用的数据库架构,以支持多国业务增长
- 改进现有的数据库基础设施,包括性能调优、容量规划、备份/恢复策略和灾难恢复流程
- 使用 Go/Python 开发自动化工具和内部平台,简化数据库操作,减少人工干预,降低人为错误
- 监控数据库健康状况和性能指标,基于 SLO 耗尽率建立告警机制,主动识别潜在问题
- 对数据库运维负责,包括模式变更、数据迁移和生产部署
- 与开发团队合作,提供数据库设计咨询、查询优化和故障排查支持
- 定期进行数据库安全审计,实施安全最佳实践,保护敏感数据
你将带来的能力
- 4 年以上处理和管理高级 MySQL 和 MongoDB 问题的经验
- 自驱力强的数据库可靠性工程师,有在生产环境中管理大规模数据库系统的经验
查看英文原文
About the role
With tens of millions of visitors every week, we strive to be the best in the industry for our users. You’ll be responsible for taking ownership of our existing and new database infrastructure for multiple platforms and facilitating releases into new regions. We wrote our systems from scratch about 4 years ago, so you’ll be working with the latest technology and won’t have to worry about decades old technologies. There’s also freedom to research and implement new technologies into our stack as you and the team see fit.
Our Stack
- Relational Databases: AWS Aurora MySQL/PostgreSQL
- NoSQL Databases: MongoDB Atlas, AWS DocumentDB
- Caching Layer: AWS ElastiCache, Valkey
- Message Queue: Apache RocketMQ, RabbitMQ
- Search Engine: ElasticSearch, Mongo Atlas
- OLAP: Redshift
- Monitoring & Alerting: Prometheus, Grafana, Loki, Alert Manager
- Backup & Recovery: AWS Backup, Point-in-Time Recovery, Cross-Region Replication
- Database Proxy: RDS Proxy
- Infrastructure as Code: Terraform, Ansible
- CI/CD Integration: Jenkins, ArgoCD, Github Action, Helm
- Network & Security: AWS VPC, Security Groups, AWS Secrets Manager
- Programming Languages: Golang, Python
- Containerization: Kubernetes
What you'll be doing
- Work in a team of DBRE and DevOps professionals
- Design, implement, and maintain highly available database architectures to support business growth across multiple countries
- Improve existing database infrastructure, including performance tuning, capacity planning, backup/recovery strategies, and disaster recovery procedures
- Develop automation tools and internal platforms using Go/Python to streamline database operations, reduce manual interventions, and minimize human errors
- Monitor database health and performance metrics, establish alerting mechanisms based on SLO burn rates, and proactively identify potential issues
- Take ownership and responsibility for database operations including schema changes, data migrations, and production deployments
- Collaborate with development teams to provide database design consultation, query optimization, and troubleshooting support
- Conduct regular database security audits and implement security best practices to protect sensitive data
What you'll bring
- 4+ years experience tackling and managing advanced MySQL and MongoDB problems
- A self-driven Database Reliability Engineer with proven experience in managing large-scale database systems in production environments
- Have a deep understanding of database architecture, performance optimisation, high availability solutions, and database security best practices
- Following the latest industry trends in database technologies and are passionate about building reliable, scalable data infrastructure
- Have strong automation mindset and programming skills to build tools that improve database operations and reliability
- Understand SLI/SLO/SLA concepts and have experience implementing reliability engineering practices
What’s in it for you
- Sporty is a remote first company in pursuit of sustainability
- A competitive salary + individual performance based bonuses every quarter
- 28 days paid annual leave
- Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
- Referral bonuses & flash bonuses
- Top of the line equipment
- Annual company retreats to provide great internal networking opportunities
Interview Process
- Remote video screening with our Talent Acquisition Team
- Online assessment via Hackerrank
- Remote video interview with Team Members (60 Mins)
- Final discussion with the hiring manager (60 mins)
If you're interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.