远程工作雷达

站点可靠性工程师(变现)

Site Reliability Engineer (Monetization)

开发工程未标注地域
公司Xsolla
薪资未公开
工作地点Montreal
地域资格未标注地域
时区要求无特别要求
用工类型Full time
发布时间未知
数据来源Lever
前往企业招聘页投递 →

你的情况

我们正在寻找一位站点可靠性工程师(变现),他务实、以产品为导向,同时擅长编写代码和运行生产系统,加入我们基础设施部门的SRE团队。最佳候选人应能在快节奏、高度协作且极具动态的环境中茁壮成长,并热衷于从头到尾负责高流量电商领域的应用级基础设施和可靠性——从部署流水线和Kubernetes清单到SLO、容量规划和生产就绪性。

具备强大的Kubernetes、可观测性和软件工程技能至关重要,同时还需有在云环境中(GCP/GKE或类似)运维生产服务的经验,并与产品开发团队紧密合作。能够从双重视角出发——理解开发者如何交付功能以及基础设施需要什么才能保持可靠——并在设计决策早期引入可靠性视角,将是你在这个职位上取得成功的关键。

这是一个混合嵌入式职位:你仍然是SRE组织的一部分(实践、标准、值班轮换),同时功能性地嵌入到变现产品领域。你将与该领域的工程团队建立长期的工作关系,负责他们应用基础设施执行的重要部分,并共同撰写公司范围内使用的可靠性实践。

如果你热衷于让复杂的分布式系统变得枯燥地可靠,并热爱构建让全球游戏开发者获得报酬的电商和变现基础架构,我们很期待收到你的来信!

关于我们

Xsolla是一家全球商业公司,拥有强大的工具和服务,帮助开发者解决视频游戏行业的固有挑战。从独立开发到AAA级别,公司与Xsolla合作,帮助他们资助、分发、营销和变现他们的游戏。基于对视频游戏未来的确信,Xsolla致力于汇聚机会,并持续为创作者提供新的资源。总部和注册地均位于加利福尼亚州洛杉矶,Xsolla作为交易商记录方运营,并已帮助超过1500名游戏开发者触达更多玩家并在全球范围内增长业务。随着更多盈利路径和赢利方式,开发者拥有了享受游戏所需的一切。

更多信息,请访问 xsolla.com。

福利

我们热衷于

查看英文原文

ABOUT YOU

We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our Infrastructure department's SRE team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs, capacity planning, and production readiness.

Strong Kubernetes, observability, and software engineering skills are essential, along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product development teams. The ability to hold a dual perspective - understanding both how developers ship features and what infrastructure needs to stay reliable - and to bring the reliability lens into design decisions early will be key to your success in this role.

This is a hybrid embedded role: you remain part of the SRE organization (practices, standards, duty rotation) while being functionally embedded into the Monetization product domain. You'll build long-term working relationships with the domain's engineering teams, own a meaningful share of their application infrastructure execution, and co-author the reliability practices used company-wide.

If you're passionate about making complex distributed systems boringly reliable and love building the commerce and monetization backbone that lets game developers around the world get paid, we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

For more information, visit xsolla.com.

Benefits

We are passionate about fostering a supportive environment for our team, so we prioritize the physical, mental, and emotional well-being of our employees and their families through a comprehensive Benefits Program. This includes medical, dental, and vision, PTO, and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities, we ensure that our team thrives both personally and professionally. Together, we're not just building a business; we're cultivating a community that values creativity, collaboration, and the transformative power of play.

Equal Employment Opportunity Statement

Xsolla is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or any other characteristic protected by law. We consider qualified applicants with criminal histories in accordance with the Fair Chance Act.

Criminal History Consideration

For the Site Reliability Engineer (Monetization) position, we will conduct a background check that may include the following:

  • Criminal history check
  • Employment verification
  • Education verification

Relevance to Job Responsibilities

The background check is relevant to this position because of the following role responsibilities:

  • Accessing confidential company data
  • Handling infrastructure that processes sensitive financial transactions
  • Ensuring compliance with regulatory requirements

Rights Under the Fair Chance Act

Applicants are encouraged to inquire about their rights under the Fair Chance Act. If you have questions regarding our hiring practices, please contact careers@xsolla.com.

By submitting the following job application form, you consent to Xsolla processing your data for career-related inquiries and potential employment opportunities. We process your data in accordance with this Xsolla Privacy Notice for Job Applicants. Please direct any inquiries regarding your data privacy to careers@xsolla.com.

Responsibilities

  • Own the application-level infrastructure of the Monetization domain: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations
  • Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling
  • Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation
  • Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation
  • Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain
  • Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks
  • Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts
  • Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads
  • Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change
  • Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems
  • Participate in the SRE duty rotation, supporting developers across the company

Qualifications & Skills

  • 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services
  • Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP)
  • Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes)
  • Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry
  • Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
  • GCP experience (IAM, networking, managed services)
  • Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)
  • Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)
  • Practical experience with incident response, post-mortems, and driving reliability improvements from incidents
  • Strong collaboration and communication skills — this role works embedded with product development teams daily
  • Experience in payments, fintech, e-commerce, or gaming — high-traffic transactional systems

Nice to Have:

  • Kubernetes certifications
  • Google Cloud Platform certifications
  • HashiCorp certifications
本页面信息整理自 Lever,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位