远程工作雷达

中级SRE分析师

Mid-level SRE Analyst

开发工程限定地区(需当地身份)与中国几乎无重叠,需长期倒时差
公司Experian
薪资未公开
工作地点Brazil
地域资格限定地区(需当地身份)
时区要求与中国几乎无重叠,需长期倒时差
用工类型permanent
发布时间4 天前
数据来源4dayweek.io
前往 4dayweek.io 查看并投递 →
注意地域限制:该职位明确限定在 Brazil 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。
作息提示:与中国几乎无重叠,需长期倒时差。

#### 公司简介

Experian 是一家全球数据和技术公司,为世界各地的个人和企业创造机会。我们在多个市场运营,包括金融服务、医疗保健、汽车、农业业务、保险等。Experian 投资于人才和先进技术,以释放数据的潜力。我们在 32 个国家拥有 25,200 名员工。

我们的独特之处在于重视你的价值。Experian 的以人为本、包容和目标驱动的文化获得了众多奖项的认可——包括 2025 年《财富》全球最佳工作场所™(全球前 25 强)以及在 26 个国家获得的“最佳工作场所™”等。查看 Experian Life 的社交媒体或浏览我们的职业网站,了解原因。Experian 还自豪地成为平等机会和积极行动雇主。

#### 职位描述

我们正在寻找一位高度积极的 **中级站点可靠性工程师(SRE)** 加入我们的 **云、数据与 AI 平台** 团队。在此职位上,你将负责设计、运营并持续改进支持业务关键应用、数据流水线和 AI/ML 工作负载的云原生平台的可靠性、可扩展性、可观测性和性能。

你将与软件工程、数据工程、AI 工程和平台团队紧密合作,构建弹性系统,自动化运维,提升开发者体验,并在整个组织中建立可靠性最佳实践。

作为中级工程师,你将在通过自动化、可观测性、事件管理、成本优化和基础设施现代化推动运营卓越方面发挥关键作用。

**主要职责:**

- 在 AWS 上设计和运营高可用、可扩展且安全的云平台。
- 构建和维护基于 Kubernetes 的基础设施,以支持应用程序、数据和 AI 工作负载。
- 通过自动化、基础设施即代码(IaC)和自助服务功能提升平台可靠性。
- 使用 Datadog 实现和增强可观测性解决方案,包括监控、日志、追踪、警报、仪表板和 SLO 管理。
- 使用 Airflow、Amazon EMR、S3 和其他 AWS 数据服务支持和优化大规模数据处理环境。
- 与数据和 AI 团队合作,提升机器学习和人工智能平台的可靠性、可扩展性和运营成熟度。
- 领导事件响应活动,分析根本原因

查看英文原文

#### Company Description

Experian is a global data and technology company that powers opportunities for people and businesses around the world. We operate in diverse markets, such as financial services, healthcare, automotive, agribusiness, insurance, among others. Experian invests in people and new advanced technologies to unlock the power of data. We have an incredible team of 25,200 employees in 32 countries.

Our uniqueness is valuing yours. Experian's people-centric, inclusive, and purpose-driven culture is recognized by numerous awards — including World’s Best Workplaces™ 2025 (Fortune's Top 25 global) and Great Place To Work™ in 26 countries, among others. Check out Experian Life on social media or explore our careers site to understand why. Experian is also proud to be an equal opportunity and affirmative action employer.

#### Job Description

We are looking for a highly motivated **Mid-level Site Reliability Engineer (SRE)** to join our **Cloud, Data & AI Platform** team. In this role, you will be responsible for designing, operating, and continuously improving the reliability, scalability, observability, and performance of cloud-native platforms that support business-critical applications, data pipelines, and AI/ML workloads.

You will work closely with Software Engineering, Data Engineering, AI Engineering, and Platform teams to build resilient systems, automate operations, improve developer experience, and establish reliability best practices across the organization.

As a mid-level engineer, you will play a key role in advancing operational excellence through automation, observability, incident management, cost optimization, and infrastructure modernization.

**Key Responsibilities:**

- Design and operate highly available, scalable, and secure cloud platforms on AWS.
- Build and maintain Kubernetes-based infrastructure to support applications, data, and AI workloads.
- Improve platform reliability through automation, Infrastructure as Code (IaC), and self-service capabilities.
- Implement and enhance observability solutions using Datadog, including monitoring, logs, tracing, alerts, dashboards, and SLO management.
- Support and optimize large-scale data processing environments using Airflow, Amazon EMR, S3, and other AWS data services.
- Partner with Data and AI teams to improve the reliability, scalability, and operational maturity of Machine Learning and Artificial Intelligence platforms.
- Lead incident response activities, root cause analysis, and post-incident reviews, promoting continuous improvement.
- Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Enhance deployment processes, CI/CD pipelines, and release reliability.
- Optimize cloud infrastructure utilization, performance, and costs.
- Mentor team members and promote SRE best practices across the engineering organization.

**What defines success in this role**

- Increased platform availability and reliability.
- Improved observability and reduced incident resolution time.
- Greater automation and reduction of manual, repetitive operational efforts.
- Reliable, scalable, and cost-efficient Data and AI platforms.
- Strong collaboration with Engineering teams to deliver resilient production systems.

#### Qualifications

**Required Qualifications**

- Higher education in progress or completed.
- Solid experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps roles.
- Strong hands-on experience with AWS services and cloud-native architectures.
- Deep knowledge of Kubernetes and containerized workloads in production environments.
- Experience managing and troubleshooting large-scale distributed systems.
- Solid experience with observability platforms, preferably Datadog.
- Experience supporting platforms and data flows using technologies such as Airflow, EMR, Spark, and S3.
- Expertise in Infrastructure as Code (IaC) using Terraform or similar tools.
- Experience building and maintaining CI/CD pipelines and platform automation.
- Solid knowledge of Linux, networking, and system performance troubleshooting.
- Proficiency in scripting and automation using Python, Bash, or similar languages.

**Desirable Qualifications**

- Experience supporting large-scale cloud-native platforms in AWS environments.
- Experience with Kubernetes platform operations and cluster lifecycle management.
- Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets, and operational excellence practices.
- Experience implementing observability solutions using tools such as Datadog, Prometheus, Grafana, OpenTelemetry, or similar technologies.
- Familiarity with data processing and workflow orchestration platforms, such as Airflow, Spark, or EMR.
- Experience with Infrastructure as Code (IaC) and platform automation practices.
- AWS, Kubernetes, Terraform, or Datadog certifications.
- Experience working in large-scale, highly available, or mission-critical corporate environments.
- Intermediate technical English.

#### Additional Information

This is an affirmative position for women, as part of our commitment to advancing gender equity in the workplace.

Serasa Experian invests in female leadership and has a partnership with the TODAS Group. Additionally, we adhere to the 'Elas Lideram 2030' movement and have joined UN Women to reduce gender inequality by 2030!!

We also have the Women in Experian group, which seeks to improve gender equity for women by creating development opportunities, different perspectives, skills, among others.

Come be part of this transformation!

If you do not identify with this group, we invite you to explore other opportunities on our careers page. We have several vacancies that may connect with your profile and interests.

**#LI-HOME**

Experian Careers - Creating a better tomorrow together

[Find out what its like to work for Experian by clicking here](https://www.experian.com/careers/)

本页面信息整理自 4dayweek.io,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

企业业务发展代表

ExperianUnited States$37,500 - $62,500/年permanent今天
市场运营职能支持限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

高级总监,需求生成

ExperianUnited States$176,036 - $316,865/年permanent昨天
市场运营限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

专家合规风险顾问

ExperianUnited States$80,237 - $139,077/年permanent昨天
职能支持限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

杰出平台解决方案架构师

ExperianUnited States$208,515 - $375,327/年permanent5 天前
AI开发工程限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

← 返回全部职位