工程团队负责人(站点可靠性工程)
Engineering Team Leader (Site Reliability Engineering)
我们正在打造XTB——一家全球投资公司,提供创新的技术解决方案,让客户以多种方式有效管理财务。这一切都通过一个直观的XTB应用实现,该应用已在全球超过一百万用户中使用!
我们是一家经过认证的优秀工作场所公司。
我们正在寻找一位工程团队负责人,负责推动站点可靠性工程(SRE)团队的开发和成长。在此职位中,您将有机会塑造SRE实践的技术和运营方向,领导弹性策略,并在定义和交付确保XTB系统可靠性和可扩展性的解决方案中发挥关键作用,这些系统服务于不断壮大的组织中的数百万客户。
职责
- 团队领导:打造并发展一支高绩效的站点可靠性工程团队,营造技术卓越、责任意识和持续改进的环境。
- 可靠性策略:与基础设施和开发团队紧密合作,制定并推动SRE平台策略,确保与组织业务目标保持一致。监督整个组织的可靠性策略,确保架构的一致性、可扩展性以及行业标准SRE最佳实践的采用。
- 事件管理:负责并监督全组织范围内的7x24小时值班和事件管理流程。管理事件工具,建立并维护操作流程,确保符合法规要求,监督报告,并推动事件响应策略的持续改进。
- 数据驱动管理:为团队绩效定义并跟踪可衡量的目标(KPI)。利用数据推动改进,构建能清晰展示运营健康状况和团队生产力的管理指标。
- 可观测性工程:负责组织范围内可观测性生态系统的设计、开发和演进。领导标准化遥测的实施策略,包括结构化日志、分布式追踪和智能采样。
要求
- 专业背景:多年在SRE、基础设施或DevOps角色中管理高规模、分布式环境的经验。
- 领导经验:在正式管理岗位上有成功记录,能够领导、指导和发展高绩效的SRE或DevOps工程团队。
- 基础设施与可靠性:在构建可靠系统方面有丰富经验,熟悉现代运维实践。
查看英文原文
We are building XTB – a global investment company offering innovative technological solutions that allow our clients to effectively manage their finances in multiple ways. All of this within a single, intuitive XTB app already used by over one million users worldwide!
We are a certified Great Place to Work company.
We are looking for an Engineering Team Leader to drive the development and growth of the Site Reliability Engineering team. In this role, you will have the opportunity to shape the technical and operational direction of SRE practices, lead a resilience strategy, and play a key role in defining and delivering solutions that ensure the reliability and scalability of XTB systems for millions of clients across a growing organization.
Responsibilities
- Team Leadership: Shape and grow a high-performing Site Reliability Engineering team, fostering an environment of technical excellence, ownership, and continuous improvement.
- Reliability Strategy: Define and drive the SRE platform strategy in close collaboration with infrastructure and development teams, ensuring alignment with organizational business objectives. Oversee reliability strategy across the organization, ensuring consistent architectural alignment, scalability, and the adoption of industry-standard SRE best practices.
- Incident Management: Own and oversee organization-wide 24/7 on-call and incident management processes. Manage incident tooling, establish and maintain operational procedures, ensure compliance with regulations, oversee reporting, and drive continuous improvement of incident response strategy.
- Data-Driven Management: Define and track measurable objectives (KPIs) for team performance. Leverage data to drive improvements and build management metrics that provide clear visibility into operational health and team productivity.
- Observability Engineering: Oversee the design, development, and evolution of the organization-wide observability ecosystem. Lead the strategy for implementing standardized telemetry, including structured logging, distributed tracing, and intelligent sampling.
Requirements
- Professional Background: Several years of experience in SRE, Infrastructure, or DevOps roles managing high-scale, distributed environments.
- Leadership Experience: Proven track record in a formal management role, leading, mentoring, and developing high-performing SRE or DevOps engineering teams.
- Infrastructure & Reliability: Extensive experience building and maintaining scalable, reliable, and observable infrastructure systems (Azure, Kubernetes, On-prem).
- Reliability Strategy: Demonstrated ability to deliver end-to-end reliability strategies, drive architectural improvements, and manage large-scale technical projects from design to production.
- Cross-Functional Collaboration: Proven ability to drive cultural change, act as a strategic partner to product engineering teams, and work effectively with distributed/remote teams.
Leadership & Management
- Execution: Break down complex projects into actionable tasks; drive incremental value.
- Growth: Mentor and support the professional growth of team members.
- Collaboration: Facilitate workshops; build technical community; resolve conflicts effectively.
- Operational Excellence: Drive operational excellence and reliability culture within the team; lead incident management and champion post-mortems.
- Strategy: Proactively manage technical debt; align team output with organizational goals.
Technical skills we expect
- Programming & Scripting: Strong Python skills for building scalable automation, internal tools, and scripts.
- Cloud & Orchestration: Expertise in managing Kubernetes, configuration management with Ansible, and designing resilient infrastructure on Azure and on-prem.
- Observability Engineering: Deep proficiency in building standardized telemetry systems. Mastery of tools like Prometheus, Grafana, OTEL, ELK, Tempo, Thanos, and similar.
- AI & Automation: Using AI/ML for AIOps, anomaly detection, log analysis, and optimizing reliability workflows.
Nice to have
- Experience with commercial APM platforms (e.g., Datadog, Splunk, New Relic) and chaos engineering tooling.
- Proficiency with cloud cost management and FinOps principles.
What we offer
- Real influence on the development of the company and the product.
- Work in an experienced team that is happy to share its knowledge.
- A clear vision of development thanks to regular feedback and clear career paths.
- Regular team-building meetings.
Benefits
- A training budget for courses and conferences that interest you.
- An extra day off on your birthday.
- An extra day off for parents.
- Equipment tailored to your needs.
- Private medical care and group insurance.
- Access to an e-learning platform for learning English and a benefits platform.
- Access to a wellbeing platform and the opportunity to take advantage of workshops and private therapy sessions.
- Remote work, from the office in Warsaw or from a coworking space in your city.
Originally posted on Himalayas