远程工作雷达

高级站点可靠性工程师

Senior Site Reliability Engineer

开发工程未标注地域
公司Playson
薪资未公开
工作地点Ukraine
地域资格未标注地域
时区要求日间重叠约 3 小时,需偶尔早起或晚睡
用工类型permanent
发布时间15 天前
数据来源4dayweek.io
前往 4dayweek.io 查看并投递 →

**职位描述**

我们正在寻找一位**高级站点可靠性工程师**加入我们的基础设施小组——这是一个精干且资深的团队,拥有很高的责任感和更高的期望。这是一项核心的高负载系统中的深入实践工作,你将直接负责在快节奏环境中维护系统的可靠性、性能和稳定性。

你将处理实时生产环境中的挑战,处理事件,管理警报,并参与关键的轮班值班。这个职位需要在高压环境下做出果断决策,并具备积极主动的心态,以持续改进大规模系统的运行。

如果你能在高负载环境中茁壮成长,喜欢解决复杂的生产问题,并希望对数百万用户使用的系统产生直接影响,那么这里就是你的理想之地。

**主要职责**

- 通过主动监控平台健康状况、管理警报并实时响应事件来确保系统可靠性

- 参与24/7的轮班值班,全面负责高流量(5–7k RPS)环境下的生产稳定性

- 调查事件,进行根本原因分析,并实施长期解决方案以防止再次发生

- 在Kubernetes(EKS)生态系统中构建并持续改进监控、警报和可观测性

- 使用Terraform、Helm和GitOps工具(Flux/ArgoCD)部署、管理和优化基础设施

- 推动自动化,积极提升系统弹性,减少人工干预和重复性问题

- 维护和演进CI/CD流水线和基础设施即代码实践

- 与工程团队紧密合作,支持部署并在生产环境中最小化用户影响

- 引入和集成新工具和技术,以增强可扩展性、可靠性和性能

- 处理特定环境的请求,确保在持续高负载下平台日常操作的顺畅

**要求**

- 在高负载环境中具有使用Kubernetes(部署、扩展、故障排除)的实际经验

- 具有使用GitOps工具如FluxCD或ArgoCD的经验

- 在生产系统中具有事件响应、根本原因分析和事后总结的实操经验

- 熟悉AWS、Terraform、Docker和CI/CD流水线

- 具有使用监控和可观测性工具如Datadog、Prometheus、Grafana以及日志堆栈如ELK或CloudWatch的经验

- 对网络概念有深入理解

查看英文原文

**About the Role**

We’re looking for a **Senior Site Reliability Engineer** to join our Infrastructure Squad - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment.

You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale.

If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you.

**Key Responsibilities**

- Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time

- Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment

- Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence

- Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem

- Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)

- Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues

- Maintain and evolve CI/CD pipelines and infrastructure-as-code practices

- Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment

- Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance

- Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load

**Requirements**

- Strong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environments

- Experience with GitOps tools such as FluxCD or ArgoCD

- Proven experience in incident response, root cause analysis, and postmortems in production systems

- Solid experience with AWS, Terraform, Docker, and CI/CD pipelines

- Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch

- Strong understanding of networking concepts and protocols

- Proficiency in at least one scripting language (e.g. Python, Go, Node.js)

- Experience working with version control systems (Git)

- Familiarity with incident management tools like PagerDuty, Opsgenie, or similar

- Ability to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountability

- Proactive, resilient mindset with a focus on continuous improvement and system stability

**What We Offer**

- Competitive Salary

- Quarterly Bonuses

- Unlimited Paid Time Off

- Unlimited Paid Sick Leave

- Remote & Flexible Working

- Private Medical Insurance

- Financial Support for Life Events

- Professional Development Budget

- International Exposure

- Regular Company Events

_*Benefits may vary depending on location and contractual agreement_

**Recruitment Process**

1. HR Interview (30-45 min)

2. Technical interview (90 min)

4. Final Interview with C-level (60 min)

_By submitting your application, you acknowledge that your personal data will be processed in accordance with our_ [_Privacy Policy_](https://playson.com/privacy) _._

本页面信息整理自 4dayweek.io,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

内容经理

Playson远程permanent13 天前
市场运营全球可投(据职位描述推断)

技术支持工程师

Playson远程permanent15 天前
开发工程市场运营全球可投(据职位描述推断)

市场设计师

Playson远程permanent15 天前
设计市场运营全球可投(据职位描述推断)

← 返回全部职位