周末系统可靠性工程师
Weekend Site Reliability Engineer
开发工程全球可投
公司Sporty Group
薪资未公开
工作地点Global - Remote
地域资格全球可投
时区要求无特别要求
用工类型未标注
发布时间2026-01-06
数据来源Greenhouse
全球可投:该职位未限制候选人所在地区。仍需注意薪资可能按地区折算,以及实际签约方式(正式雇佣 / 独立合同)。
你将负责的工作
- 与 DevOps 和 DBA 专业团队合作;作为周末 SRE,覆盖周六、周日和周一(总共5天,休息日有灵活性)
- 改进我们部署国家的现有基础设施和流程,同时优化未来部署到新国家的流程
- 持续提升 Kubernetes 平台的稳定性和效率,重点在于优化资源利用率、降低成本,并通过 GitOps 优先实践简化环境配置
- 通过自动扩展、警报管道和 Grafana 仪表盘监控和维护云基础设施,涵盖指标、日志、追踪和真实用户监控(RUM)
- 负责周末的值班运营,对生产事件进行分类和响应,执行根本原因分析,并推动事后复盘
- 设计和管理警报管道,确保警报信号的可操作性,注意防止警报疲劳、瀑布式警报和通知泛滥
- 定义并维护关键服务的 SLIs 和 SLOs,并利用它们推动可靠性改进和值班优先级排序
- 对我们的云运营活动负责并承担责任
- 与外部安全机构联络进行年度审计,同时执行我们自己的内部安全检查
- 协助重新配置现有架构,以实现快速部署到新国家
- 指导经验较少的团队成员
你将带来的能力
- 3年以上 DevOps/平台工程经验
- 必须位于欧洲、亚洲或拉美地区
- 有独立领导项目规划和部署的经验
- 熟悉云平台,尤其是 AWS,具备扎实的知识来利用云资源满足其他团队和生产的需求
- 对 Kubernetes 和容器编排有深入理解,具有 EKS 和 GitOps 工具如 ArgoCD 和 Helm 的经验者尤为看重
- 具有 Infrastructure-as-Code 经验,特别是 Terraform
- 熟练使用 Bash、Python 或 Golang 进行脚本编写和自动化,有 Rust 经验者优先
- 具有覆盖指标、日志、分布式追踪和性能分析的可观测性工具栈的实际经验,例如 Prometheus、Loki、Tempo、Pyroscope 和 OpenTelemetry
- 具有真实用户监控(RUM)经验,熟悉 Grafana Faro 或 OpenTelemetry SDK 集成者优先
- 有值班经验
查看英文原文
What you’ll be doing
- Work with a team of DevOps and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SRE
- Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
- Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
- Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
- Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
- Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
- Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
- Take ownership and responsibility for our cloud operation activities
- Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
- Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
- Mentoring less experienced team members
What you’ll bring
- 3+ years DevOps / platform engineering experience
- Must be based in Europe or Asia or LatAM
- Experience independently leading the planning and deployment of a project
- Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
- Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
- Experience with Infrastructure-as-Code, particularly Terraform
- Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
- Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
- Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
- Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
- Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
- Experience defining SLIs and SLOs and using them to inform reliability work
- Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
- Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
- Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
- A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
- Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous
Our stack
- Languages: Java / Spring Boot, Node.js, Python, JavaScript
- Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
- Cache: ElastiCache, Redis, Valkey
- Messaging: Apache RocketMQ, AutoMQ, Kafka
- Networking & Proxy: Nginx, Kong, Cilium, eBPF
- Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
- Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
- CI/CD: Jenkins, GitHub Actions
- Metrics: Prometheus, Mimir, Grafana, Alertmanager
- Logs: Loki, Vector
- Traces: Tempo, OpenTelemetry, Alloy
- Profiling: Pyroscope
- RUM: Grafana Faro, OpenTelemetry SDK
- Infrastructure as Code: Terraform
- CDN & Edge: Cloudflare, AWS CloudFront
- AWS CloudWatch
What’s in it for you
- Sporty is a remote first company in pursuit of sustainability
- A competitive salary + individual performance based bonuses every quarter
- 28 days paid annual leave
- Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
- Referral bonuses & flash bonuses
- Top of the line equipment
- Annual company retreats to provide great internal networking opportunities
Interview process
- Remote video screening with our Talent Acquisition Team
- Online assessment via Hackerrank
- Remote video interview with 3 x Team Members (45 mins each, not separate days)
If you're interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.
#LI-remote
本页面信息整理自 Greenhouse,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。
本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。