站点可靠性工程师
Site Reliability Engineer
XTB是一家来自金融行业的全球公司,专注于金融工具的在线交易。我们是波兰最大的FinTech公司,在中欧和东欧地区处于领先地位,业务范围涵盖多个国家,包括亚洲和南美洲。在XTB,我们注重员工的发展,为他们提供在各个领域获取知识和技能的机会,以及众多培训和发展计划。如果您正在寻找挑战,并希望在国际商业环境中获得宝贵的经验,XTB就是您的理想选择。
我们是一家经过认证的卓越工作场所公司。
我们正在寻找一名站点可靠性工程师,负责定义并推动XTB系统在数百万客户规模下的可靠性。在此职位中,您将加强SRE实践,并通过高影响力的可观测性,塑造我们整个技术栈的弹性,确保我们的系统保持稳健和可扩展。
职责
- 可观测性平台工程:开发标准化的可观测性生态系统。实施有意识的遥测模型,重点关注结构化事件、分布式追踪和智能采样策略,为系统行为提供深入且可操作的洞察。
- 可靠性赋能:作为产品工程团队的战略合作伙伴,为其提供平台、标准和数据,以实现服务可靠性。使用错误预算和警报作为主要语言,平衡功能速度与稳定性。
- 主动弹性与保护:增强检测能力,以便在问题影响客户之前识别它们。利用预警系统和AI/ML进行自动化异常检测和智能数据分析,持续验证和增强系统弹性。
- 运维与工具:构建内部自动化和工具,简化SRE工作流程,自动化常规运维任务,并提升整个技术栈的效率。
- 事件管理与值班轮班:参与值班轮班,提供事件管理,确保快速解决事件、有效沟通,并进行事后分析,以推动持续改进。
要求
- 专业背景:在SRE、基础设施或DevOps岗位上有经验,管理高规模、分布式环境。
- 技术工程:具备Python的高级编程技能,重点在于
查看英文原文
XTB is a global company from the financial industry, focusing on online trading of financial instruments. We are the largest FinTech in Poland and a leader in Central and Eastern Europe, and the range of our operations covers several countries, including Asia and South America. At XTB, we focus on the development of our employees, giving them opportunities to gain knowledge and skills in various fields, as well as offering a number of training and development programs. If you are looking for challenges and want to gain valuable experience in an international business environment, XTB is the right place for you.
We are a certified Great Place to Work company.
We are looking for a Site Reliability Engineer to define and drive the reliability of XTB systems at the scale of millions of clients. In this role, you will strengthen SRE practices and shape the resilience of our entire technology stack through high-impact observability, ensuring our systems remain robust and scalable.
Responsibilities
- Observability Platform Engineering: Develop a standardized observability ecosystem. Implement a conscious telemetry model focusing on structured events, distributed tracing, and intelligent sampling strategies - that provides deep, actionable insights into system behavior.
- Reliability Enablement: Act as a strategic partner to product engineering teams, providing the platform, standards, and data they need to own service reliability. Use error budgets and alerting as the primary language for balancing feature velocity with stability.
- Proactive Resilience & Protection: Enhance detection capabilities to identify issues before they impact the customer. Leverage early-warning systems and AI/ML for automated anomaly detection and intelligent data analysis to continuously verify and strengthen system resilience.
- Operations & Tooling: Build internal automation and tooling that streamlines SRE workflows, automates routine operational tasks, and enhances efficiency across the technology stack.
- Incident Management & On-Call Rotation: Participate in an on-call rotation to provide incident management, ensuring rapid incident resolution, effective communication, and post-incident analysis to drive continuous improvement.
Requirements
- Professional Background: Professional experience in SRE, Infrastructure, or DevOps roles managing high-scale, distributed environments.
- Technical Engineering: Advanced programming skills in Python, with a strong focus on building scalable automation, internal tooling, and robust scripts.
- Cloud & Orchestration: Hands-on expertise in managing production-grade Kubernetes environments, configuration management tools like Ansible, and designing resilient infrastructure architectures within Azure Kubernetes Service and on-prem environments.
- Observability Engineering: Proficiency in building standardized telemetry ecosystems. You have mastered self-hosted opensource tools for observability data collection, storage and visualization, like Prometheus, Grafana, ELK Stack, Tempo, Thanos, Jaeger and similar.
- Operational & Soft Skills: Ability to drive incident management, conduct thorough post-incident analysis, and foster a culture of reliability and shared ownership.
Nice to have
- Experience with commercial APM platforms (e.g., Datadog, Splunk, New Relic) and chaos engineering tooling.
- Experience with cloud cost management and FinOps principles.
- Experience defining and tracking SRE metrics (SLI/SLOs) and managing error budgets to drive reliability.
- Experience with AI/ML techniques for SRE tasks, such as AIOps, automated anomaly detection, log analysis, and optimizing reliability workflows.
- Experience in building and managing strategies to proactively manage technical debt and align team output with organizational goals.
What we offer
- Real influence on the development of the company and the product.
- Work in an experienced team that is happy to share its knowledge.
- A clear vision of development thanks to regular feedback and clear career paths.
- Regular team-building meetings.
Benefits
- A training budget for courses and conferences that interest you.
- An extra day off on your birthday.
- An extra day off for parents.
- Equipment tailored to your needs.
- Private medical care and group insurance.
- Access to an e-learning platform for learning English and a benefits platform.
- Access to a wellbeing platform and the opportunity to take advantage of workshops and private therapy sessions.
- Remote work, from the office in Warsaw or from a coworking space in your city.
Originally posted on Himalayas