远程工作雷达

技术运维负责人(TSO 负责人),德国

Technical Service Operations Lead (TSO Lead), Germany

职能支持限定地区(需当地身份)
公司Xsolla
薪资未公开
工作地点Berlin, Germany
地域资格限定地区(需当地身份)
时区要求无特别要求
用工类型Full time
发布时间未知
数据来源Lever
前往企业招聘页投递 →
注意地域限制:该职位明确限定在 Berlin, Germany 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

关于您

我们正在寻找一位技术运维主管(TSO主管),加入我们的全球技术运维(GTO)团队。该职位需要具备运营导向、协作能力、分析能力和出色的沟通能力。最佳候选人应能在快节奏、高度协作和极具动态的环境中茁壮成长,愿意与跨职能团队合作协调事件响应,识别生产问题中的趋势和模式,改进事件期间与合作伙伴的沟通方式,并通过事件后回顾推动持续改进。

具备强大的事件管理经验、ITIL知识以及可观测性/监控专业技能是基本要求,同时需要在支持高可用平台的游戏行业技术运维、SRE或NOC环境中具有相关经验。能够清晰有效地用英语(书面和口头)与技术及高管受众沟通将是此职位成功的关键。

如果您热衷于推动大规模的运营卓越和平台可靠性,热爱确保游戏开发者和玩家依赖的商务和支付解决方案的可靠性和正常运行,我们期待您的加入!

关于公司

Xsolla是一家全球商务公司,提供强大的工具和服务,帮助开发者解决视频游戏行业的固有挑战。从独立开发到AAA级别,公司与Xsolla合作,帮助他们资助、分发、营销和变现他们的游戏。基于对视频游戏未来发展的信念,Xsolla致力于汇聚机会,不断为创作者提供新的资源。总部位于美国加利福尼亚州洛杉矶,Xsolla作为交易商负责运营,已帮助超过1500+游戏开发者在全球范围内触达更多玩家并增长业务。随着更多盈利路径和获胜方式的出现,开发者可以享受一切所需来尽情玩游戏。

更多信息,请访问 xsolla.com。

福利

我们热衷于为团队营造一个支持性的环境,因此通过全面的福利计划优先关注员工的身心健康和情感福祉。这包括无限量的灵活休假、健身房会员资格、每月交通票以及为每位员工定制的职业发展路线。通过投资职业发展,提供培训和教育机会,我们确保员工持续成长。

查看英文原文

ABOUT YOU

We are looking for a Technical Service Operations Lead (TSO Lead) who is operationally driven, collaborative, analytical, and a strong communicator to join our Global Technical Operations (GTO) team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to help coordinate incident response alongside cross-functional teams, identify trends and patterns in production issues, improve how we communicate with partners during incidents, and drive continuous improvement through post-incident reviews.

Strong incident management experience, ITIL knowledge, and observability/monitoring expertise are essential, along with experience in technical operations, SRE, or NOC environments supporting high-availability platforms in the gaming industry. The ability to communicate clearly and effectively in English — both written and verbal — across technical and executive audiences will be key to your success in this role.

If you’re passionate about driving operational excellence and platform reliability at scale and love ensuring the reliability and uptime of commerce and payment solutions that game developers and players depend on, we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

For more information, visit xsolla.com.

Benefits

We are passionate about fostering a supportive environment for our team, so we prioritize the physical, mental, and emotional well-being of our employees through a comprehensive Benefits Program. This includes unlimited Flexible Time Off, Gym membership, monthly train ticket and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities, we ensure that our team thrives both personally and professionally.

Together, we’re not just building a business; we’re cultivating a community that values creativity, collaboration, and the transformative power of play.

The duties and responsibilities of this position may evolve over time to support the organization’s goals and individual growth. This job description is intended to outline the general nature and level of work being performed and is not intended to be an exhaustive list of all duties, responsibilities, and qualifications required.

By submitting your application, you consent to Xsolla conducting background checks, where permitted by law, after the final interview stage. All checks will comply with local regulations, and your information will be handled confidentially.

Xsolla takes your privacy seriously and will not sell or externally distribute any personal data received during the hiring process. In accordance with applicable data protection laws, Xsolla is committed to protecting your personal information and respecting your privacy.

For any inquiries related to data privacy, please contact: careers@xsolla.com

Explore more opportunities at: https://xsolla.com/careers

Responsibilities:

  • Serve as Incident Commander for major incidents — coordinating cross-functional response teams, driving investigation, making escalation decisions, and ensuring incidents are resolved within SLA targets.
  • Own all incident communications: draft and send clear, timely updates to senior leadership, Customer Success, and partner/customer contacts throughout the incident lifecycle, and manage customer-facing status page updates (status.xsolla.com).
  • Facilitate blameless Post-Incident Reviews (PIRs) for major incidents — leading root cause identification, assigning corrective actions with clear owners and deadlines, and tracking them to closure.
  • During non-incident periods, proactively analyze incident trends, recurring issues, and production bugs — identify patterns, create Problem tickets, and report findings and recommendations to product and engineering teams on a regular cadence.
  • Enforce the incident management framework across the organization, including the severity model, priority matrix, SLA targets, escalation procedures, and deployment readiness gates.
  • Oversee and mentor the Operations Engineer on your shift — coaching on triage, investigation, runbook execution, and documentation quality while conducting regular knowledge transfer sessions to build depth across the service portfolio.
  • Produce shift handoff reports and deliver regular operational reporting: incident trends, KPI performance (MTTD, MTTA, MTTR), SLA adherence, proactive detection rate, and repeat incident analysis.
  • Audit service catalogue completeness on a regular cadence and govern JIRA Service Management workflows for incident, PIR, and problem management.
  • Cover for the Operations Engineer role during absences, breaks, or surge incidents. Participate in weekend on-call rotation for major incidents.

Qualifications:

  • Previous experience working at a gaming company is required — you understand the pace, player expectations, live operations dynamics, and the operational demands of the gaming industry.
  • 6+ years of experience in incident management, SRE, NOC leadership, or technical operations in a production environment supporting high-availability, high-transaction systems.
  • Proven incident management experience — coordinating multi-team response, making real-time escalation decisions, and communicating with executive stakeholders under pressure.
  • Excellent written and verbal communication skills in English — ability to draft clear, concise executive updates at 3 AM under pressure, facilitate blameless PIRs, present operational metrics to senior leadership, and communicate incident status to customers and partners with clarity and professionalism.
  • Strong ITIL foundation — understanding of incident, problem, and change management lifecycles with practical experience implementing or operating ITIL-aligned workflows.
  • Technical depth across the observability stack — ability to read and interpret logs, traces, and metrics in Datadog (or equivalent: Grafana, Splunk, New Relic). Understanding of APM, SLOs, error budgets, burn-rate alerting, and synthetic monitoring.
  • Hands-on experience with incident tooling: Datadog, PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence.
  • Analytical mindset — ability to identify trends, patterns, and recurring issues from incident data and translate them into actionable recommendations for product and engineering teams.
  • Experience with SLA/SLO-driven operations where MTTD, MTTA, and MTTR are measured, reported, and improved.
  • Comfort with 24x7 shift-based operations as part of a follow-the-sun model with handoff overlaps. Weekend on-call (rotating) for critical severities is required.

Nice to Have:

  • Experience with customer/partner-facing incident communications and status page management.
  • Experience with or strong interest in AI/ML-assisted operations: anomaly detection, alert correlation, predictive alerting, automated remediation, or self-healing automation.
  • JIRA Service Management administration experience: workflows, SLA timers, automation rules, queues, and permissions.
  • Familiarity with Datadog Service Catalog, scorecards, and SLOs — especially burn-rate alerts and multi-window SLOs.
  • Experience building an operations function from scratch — defining processes, writing runbooks, establishing governance cadences.
  • Background in Kubernetes, cloud infrastructure (GCP preferred), microservices architecture, or distributed systems.
  • ITIL certification (Foundation or higher) is a plus but not required.
本页面信息整理自 Lever,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位