高级站点可靠性工程师(SRE)(西班牙)
Senior Site Reliability Engineer (SRE) (Spain)
高级站点可靠性工程师(SRE)
我们正在寻找一位技能娴熟且充满热情的高级站点可靠性工程师加入我们的工程赋能团队。这是在一项大型、复杂且影响深远的项目中至关重要的角色,该项目专注于解构我们的单体架构,升级我们的技术栈,并将质量和韧性嵌入到开发生命周期的每个阶段。
您将在塑造我们未来平台的过程中发挥关键作用,推动运营卓越,并培养持续改进的文化。
您将负责:
作为我们工程赋能团队的高级SRE工程师,您将:
- 架构与实现可靠性:在Azure上设计、构建和维护高可扩展性、高弹性和高性能的系统,重点在于我们的Java、Kafka和Couchbase技术栈。
- 推动现代化:作为团队的核心成员,亲自动手推进Micronaut的采用,标准化应用模板,并向托管云服务迁移。
- 提升运营卓越:制定并实施提升系统可观测性的策略(标准化日志、指标、追踪)、告警和值班流程。
- 自动化一切:在软件开发生命周期(SDLC)中推动自动化,从CI/CD流水线到基础设施配置,重点是加快交付速度并降低部署风险。
- 事故管理与学习:参与我们成熟的无责事故回顾流程,识别根本原因并实施预防措施以减少事故时间。
- 工具与标准:开发、维护并推动跨工程团队的共享标准化SRE工具和最佳实践,包括容器化(如Docker、Azure上的Kubernetes)、基础设施即代码(如Terraform)和配置管理。
- 指导与协作:为初级工程师提供技术领导力和指导,推动整个工程组织中的SRE原则和运营卓越文化。
- 战略输入:为我们的SRE和平台计划的整体技术战略和路线图做出贡献,确保与业务目标一致。
您将带来:
· 深厚的SRE专业知识:有高级站点可靠性工程师或类似职位的丰富经验,对SRE原则(错误预算、SLO/SLI、减少重复劳动)有深入理解。
- Azure云熟练度:在Azure上有丰富的实际经验,熟悉相关技术和服务。
查看英文原文
Senior Site Reliability Engineer (SRE)
We are seeking a highly skilled and passionate Senior Site Reliability Engineer to join our Engineering Enablement team. This is a critical role within a large, complex, and high-impact initiative focused on deconstructing our monolithic architecture, revitalising our technology stack, and embedding quality and resilience into every stage of our development lifecycle.
You will play a pivotal role in shaping our future-state platform, driving operational excellence, and fostering a culture of continuous improvement.
What You'll Do:
As a Senior SRE Engineer in our Engineering Enablement team, you will:
- Architect and Implement Reliability: Design, build, and maintain highly scalable, resilient, and performant systems on Azure, focusing on our Java, Kafka, and Couchbase stack.
- Drive Modernisation: Work hands-on as part of the team spearheading the adoption of Micronaut, standardising application templates, and transitioning to managed cloud services.
- Enhance Operational Excellence: Develop and implement strategies for improving system observability (standardised logging, metrics, tracing), alerting, and on-call practices.
- Automate Everything: Champion automation across the software development lifecycle (SDLC), from CI/CD pipelines to infrastructure provisioning, focusing on accelerating delivery and de-risking deployments.
- Incident Management & Learning: Contribute to our mature, blameless post- incident review process, identifying root causes and implementing preventative measures to reduce incident hours.
- Tooling & Standards: Develop, maintain, and drive the adoption of shared, standardised SRE tooling and best practices across engineering teams, including containerisation (e.g., Docker, Kubernetes on Azure), infrastructure as code (e.g.,Terraform), and configuration management.
- Mentorship & Collaboration: Provide technical leadership and mentorship to junior engineers, fostering a culture of SRE principles and operational excellence across the wider engineering organisation.
- Strategic Input: Contribute to the overall technical strategy and roadmap for our SRE and platform initiatives, ensuring alignment with business objectives.
What You'll Bring:
- Deep SRE Expertise: Proven experience as a Senior Site Reliability Engineer or a similar role, with a strong understanding of SRE principles (error budgets, SLOs/SLIs, toil reduction).
- Azure Cloud Proficiency: Extensive hands-on experience designing, deploying, and operating highly available and scalable applications on Microsoft Azure.
- Azure Kubernetes Service (AKS) Expertise: Mandatory extensive hands-on experience with AKS for container orchestration, including deployment, scaling, monitoring, and troubleshooting.
- Java Ecosystem Mastery: Expert-level proficiency with Java, including experience with modern frameworks (ideally Micronaut, Spring Boot, or similar) and JVM performance tuning.
- Distributed Systems Knowledge: Solid understanding and practical experience with distributed systems, microservices architecture, and associated challenges (e.g., consistency, fault tolerance).
- Messaging & Database Expertise: Hands-on experience with an event streaming platform (ideally Kafka) and NoSQL data storage (ideally Couchbase), including operational best practices.
- Automation First Mindset: Strong scripting skills (e.g., Python, Bash) and experience with Infrastructure as Code tools (e.g., Terraform, ARM templates) and CI/CD pipelines (e.g., Azure DevOps, Jenkins).
- Observability Tools: Experience with monitoring, logging, and alerting tools (e.g., Azure Monitor, Prometheus, Grafana, ELK Stack, Splunk).
- Problem-Solving Acumen: Exceptional analytical and troubleshooting skills, with a methodical approach to diagnosing and resolving complex production issues.
- Communication & Collaboration: Excellent communication skills, with the ability to articulate complex technical concepts to diverse audiences and collaborate effectively with cross-functional teams.
- Continuous Improvement: A proactive and innovative mindset, always seeking ways to improve systems, processes, and team efficiency.
Some of the benefits you’ll enjoy working with us:
- The chance to join an organization with triple-digit growth that is changing the paradigm on how software products are built.
- The opportunity to form part of an amazing, multicultural community of tech expert
- A highly competitive compensation package.
• Medical insurance.
• English lessons.
Come and join our #ParserCommunity.
Follow us on Linkedin
Originally posted on Himalayas