高级站点可靠性工程师
Senior Site Reliability Engineer
**职位描述**
我们正在寻找一位**高级站点可靠性工程师**加入我们的基础设施小组——这是一个精干且资深的团队,拥有很高的责任感和更高的期望。这是一项核心的高负载系统中的深入实践工作,你将直接负责在快节奏环境中维护系统的可靠性、性能和稳定性。
你将处理实时生产环境中的挑战,处理事件,管理警报,并参与关键的轮班值班。这个职位需要在高压环境下做出果断决策,并具备积极主动的心态,以持续改进大规模系统的运行。
如果你能在高负载环境中茁壮成长,喜欢解决复杂的生产问题,并希望对数百万用户使用的系统产生直接影响,那么这里就是你的理想之地。
**主要职责**
- 通过主动监控平台健康状况、管理警报并实时响应事件来确保系统可靠性
- 参与24/7的轮班值班,全面负责高流量(5–7k RPS)环境下的生产稳定性
- 调查事件,进行根本原因分析,并实施长期解决方案以防止再次发生
- 在Kubernetes(EKS)生态系统中构建并持续改进监控、警报和可观测性
- 使用Terraform、Helm和GitOps工具(Flux/ArgoCD)部署、管理和优化基础设施
- 推动自动化,积极提升系统弹性,减少人工干预和重复性问题
- 维护和演进CI/CD流水线和基础设施即代码实践
- 与工程团队紧密合作,支持部署并在生产环境中最小化用户影响
- 引入和集成新工具和技术,以增强可扩展性、可靠性和性能
- 处理特定环境的请求,确保在持续高负载下平台日常操作的顺畅
**要求**
- 在高负载环境中具有使用Kubernetes(部署、扩展、故障排除)的实际经验
- 具有使用GitOps工具如FluxCD或ArgoCD的经验
- 在生产系统中具有事件响应、根本原因分析和事后总结的实操经验
- 熟悉AWS、Terraform、Docker和CI/CD流水线
- 具有使用监控和可观测性工具如Datadog、Prometheus、Grafana以及日志堆栈如ELK或CloudWatch的经验
- 对网络概念有深入理解
查看英文原文
**About the Role**
We’re looking for a **Senior Site Reliability Engineer** to join our Infrastructure Squad - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment.
You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale.
If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you.
**Key Responsibilities**
- Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time
- Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment
- Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence
- Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem
- Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)
- Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues
- Maintain and evolve CI/CD pipelines and infrastructure-as-code practices
- Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment
- Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance
- Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load
**Requirements**
- Strong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environments
- Experience with GitOps tools such as FluxCD or ArgoCD
- Proven experience in incident response, root cause analysis, and postmortems in production systems
- Solid experience with AWS, Terraform, Docker, and CI/CD pipelines
- Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch
- Strong understanding of networking concepts and protocols
- Proficiency in at least one scripting language (e.g. Python, Go, Node.js)
- Experience working with version control systems (Git)
- Familiarity with incident management tools like PagerDuty, Opsgenie, or similar
- Ability to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountability
- Proactive, resilient mindset with a focus on continuous improvement and system stability
**What We Offer**
- Competitive Salary
- Quarterly Bonuses
- Unlimited Paid Time Off
- Unlimited Paid Sick Leave
- Remote & Flexible Working
- Private Medical Insurance
- Financial Support for Life Events
- Professional Development Budget
- International Exposure
- Regular Company Events
_*Benefits may vary depending on location and contractual agreement_
**Recruitment Process**
1. HR Interview (30-45 min)
2. Technical interview (90 min)
4. Final Interview with C-level (60 min)
_By submitting your application, you acknowledge that your personal data will be processed in accordance with our_ [_Privacy Policy_](https://playson.com/privacy) _._