远程工作雷达

运营工程师,巴库

Operations Engineer, Baku

开发工程职能支持未标注地域
公司Xsolla
薪资未公开
工作地点Baku
地域资格未标注地域
时区要求无特别要求
用工类型Full time
发布时间未知
数据来源Lever
前往企业招聘页投递 →

你的情况

我们正在寻找一名运营工程师加入我们的全球技术运营(GTO)团队。理想的候选人应具备技术好奇心、注重细节、沟通能力强且积极主动。最佳候选人应能在快节奏、高度协作和极具动态的环境中茁壮成长,并对跨全球平台的生产问题进行监控和调查充满热情,帮助改进我们检测和响应事件的方式,分析生产数据中的趋势和模式,并在事件期间与合作伙伴和利益相关者进行更好的沟通。

强大的故障排除技能、可观测性平台经验以及脚本编写能力是必不可少的,同时需要有SRE、DevOps、生产运维或NOC环境的经验,这些环境支持高可用性平台(支付、电子商务、SaaS或游戏)。在撰写事件更新、值班交接和状态页面通讯时,能够清晰有效地用英语进行书面和口头交流将是该职位成功的关键。

如果你热衷于保持关键系统的运行并持续改进操作流程,喜欢第一时间发现并解决全球游戏开发者和玩家的问题,我们期待你的加入!

关于我们

Xsolla是一家全球性的商业公司,提供强大的工具和服务,帮助开发者解决视频游戏行业的固有挑战。从独立开发到AAA级公司,各公司都与Xsolla合作,以帮助他们为游戏融资、分发、营销和变现。基于对视频游戏未来发展的信念,Xsolla致力于汇聚机会,并不断为创作者提供新的资源。

Xsolla总部位于美国加利福尼亚州洛杉矶,作为交易商,已帮助超过1500+家游戏开发商在全球范围内触达更多玩家并增长业务。随着更多盈利途径和赢利方式的出现,开发者们拥有了享受游戏所需的一切。

更多信息,请访问 xsolla.com。

此职位的职责和责任可能会随时间推移而变化,以支持组织目标和个人成长。本职位描述旨在概述工作的总体性质和水平,而非详尽列出所有职责、责任和资格要求。通过提交申请,您表示同意。

查看英文原文

ABOUT YOU

We are looking for an Operations Engineer who is technically curious, detail-oriented, a strong communicator, and proactive to join our Global Technical Operations (GTO) team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to monitor and investigate production issues across a global platform, help improve how we detect and respond to incidents, analyze trends and patterns in production data, and contribute to better communication with partners and stakeholders during incidents.

Strong troubleshooting skills, observability platform experience, and scripting ability are essential, along with experience in SRE, DevOps, production operations, or NOC environments supporting high-availability platforms (payments, e-commerce, SaaS, or gaming). The ability to communicate clearly and effectively in English — both written and verbal — when writing incident updates, shift handoffs, and status page communications will be key to your success in this role.

If you're passionate about keeping critical systems running and continuously improving operational processes and love being the first to spot issues and the one who drives them to resolution for game developers and players worldwide, we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators.

Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

For more information, visit xsolla.com.

The duties and responsibilities of this position may evolve over time to support the organization's goals and individual growth. This job description is intended to outline the general nature and level of work being performed and is not intended to be an exhaustive list of all duties, responsibilities, and qualifications required. By submitting your application, you consent to Xsolla conducting background checks, where permitted by law, after the final interview stage. All checks will comply with local regulations, and your information will be handled confidentially. Xsolla takes your privacy seriously and will not sell or externally distribute any personal data received during the hiring process. In accordance with applicable data protection laws, Xsolla is committed to protecting your personal information and respecting your privacy.

For any inquiries related to data privacy, please contact: careers@xsolla.com

For more vacancies: Careers | Xsolla

Responsibilities

  • Serve as the primary dashboard monitor during your shift — continuously watch the GTO Operational Dashboard in Datadog, detect anomalies by correlating signals across APM, logs, metrics, synthetic tests, and Real User Monitoring, and determine whether alerts warrant an incident ticket or can be resolved through immediate investigation.
  • Triage and investigate production incidents — create incident tickets in JIRA Service Management, perform initial technical investigation using Datadog (traces, logs, infrastructure and application metrics), determine blast radius and likely root cause domain, and route to the correct team (Product SRE, Infrastructure SRE, or Engineering) using the smart routing model.
  • Own lower-severity incidents end-to-end from detection through resolution — diagnose, execute runbook procedures, and resolve without escalation where possible. Escalate promptly when an incident is unresolved within defined thresholds or requires a code-level fix.
  • Support the TSO Lead during major incidents as the technical right hand in the war room — surface real-time data (error rates, impact scope, deployment history, related alerts), maintain the incident ticket with live timeline entries and linked evidence, and execute mitigation actions as directed.
  • Draft incident communications under TSO Lead direction, including internal Slack updates, stakeholder notifications, and customer-facing status page updates (status.xsolla.com). Support clear, timely communication throughout the incident lifecycle.
  • During non-incident periods, analyze incident trends, recurring issues, and production bugs — compile data from Datadog, JIRA, and Slack, identify patterns, and contribute findings to regular reports for product and engineering teams.
  • Compile incident timelines and draft initial PIR documents for Post-Incident Review preparation. Track PIR action items post-session and flag overdue items to the TSO Lead.
  • Build and maintain operational automation (alert enrichment scripts, incident templates, Slack workflows, dashboard widgets) and contribute to runbook development — documenting new resolution procedures so they can be repeated by any Operations Engineer on any shift.
  • Conduct structured shift handoffs covering active incidents, at-risk services, upcoming deployments, and follow-up items. Participate in knowledge transfer sessions with SREs to continuously expand independent resolution capability.
  • Cover for the TSO Lead during vacations, absences, or emergencies — including severity classification, escalation decisions, stakeholder communications, and basic Incident Commander functions.
  • Publish health reports of critical apps periodically.

Qualifications and Skills

  • 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment. Experience with platforms that handle payments, e-commerce, SaaS, or gaming workloads is preferred.
  • Strong troubleshooting and investigation skills — ability to take an alert or user-reported symptom and methodically trace it through the stack: application logs, APM traces, infrastructure metrics, database queries, and network paths.
  • Hands-on experience with Datadog (or equivalent observability platform: Grafana, Splunk, New Relic, Elastic) — navigating APM, building log queries, reading infrastructure dashboards, interpreting SLO burn rates, and configuring monitors and alerts.
  • Proficiency in at least one scripting language: Python, Go, or Bash. You will write automation scripts, build operational tooling, and work with APIs.
  • Clear written and verbal communication skills in English — ability to write incident tickets, investigation notes, Slack updates, shift handoff reports, status page communications, and PIR drafts that are clear, concise, and useful to both technical and non-technical audiences.
  • Working knowledge of Kubernetes and cloud infrastructure (GCP preferred, AWS/Azure acceptable) — understanding of pods, deployments, services, ingress, node health, and how to investigate Kubernetes-related production issues.
  • Understanding of SLOs, error budgets, and burn-rate alerting — knowing what a multi-window burn-rate alert means, how error budgets deplete, and how SLO breaches translate into incident severity.
  • Experience with incident management tooling: JIRA or JIRA Service Management, PagerDuty or OpsGenie, Slack, and Confluence.
  • Experience with or strong interest in AI/ML-assisted operations: anomaly detection, alert correlation, predictive monitoring, or automated remediation.
  • Comfort with 24x7 shift-based operations as part of a follow-the-sun model with handoff overlaps. Weekend on-call (rotating) is required.

Nice to Have

  • Experience in the gaming, payments, or fintech industry — particularly environments where transaction processing, checkout flows, or player-facing services must meet strict uptime requirements.
  • Familiarity with Datadog Service Catalog, synthetic monitoring, and RUM; exposure to database operations (MySQL, PostgreSQL, Redis, Kafka); and experience with CI/CD pipelines and deployment tooling (GitLab CI, ArgoCD, Helm).
  • JIRA Service Management administration experience (workflows, automation rules, SLA timers) or ITIL Foundation certification — practical experience matters more than credentials.
本页面信息整理自 Lever,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位