云可观测性工程师
Cloud Observability Engineer
#### 公司简介
Experian 是一家全球性的数据和技术公司,为世界各地的个人和企业创造机会。我们运营于金融服务、医疗保健、汽车、农业、保险等多个市场。Experian 投资于人才和先进的新技术,以释放数据的潜力。我们在 32 个国家拥有 25,200 名员工。
我们的独特之处在于重视你的价值。Experian 的以人为本、包容和有目标导向的文化获得了众多奖项的认可——包括 2025 年世界最佳工作场所™(财富全球前 25 强)以及在 26 个国家获得的“最佳工作场所™”等。请查看 Experian 的社交媒体或浏览我们的职业网站,了解原因。Experian 还自豪地成为平等就业机会和积极行动的雇主。如果您有残疾或需要进行调整,请尽早告知我们。
#### 职位描述
云可观测性工程师的一天是关于让复杂系统变得可理解,提高信号质量,并在团队间实现更快、更智能的调试。
- 检查系统健康状况,审查警报/事件
- 处理警报
- 调查问题
- 改进可观测性工具
- 构建和改进仪表盘
- 警报优化
- 与开发团队和其他工程合作伙伴合作。
- 持续改进。
- 发布支持。
**可观测性与监控**
- 使用指标、日志和分布式追踪设计和实现可观测性框架
- 开发仪表盘、警报和可视化工具来监控系统健康状况
- 在工程团队中标准化可观测性实践(日志、遥测、追踪)
- 实现和管理原生监控工具。
- Amazon CloudWatch(指标、日志、警报、仪表盘)
- AWS X-Ray(分布式追踪)
- AWS OpenTelemetry 分发版(ADOT)
**检测与响应**
- 构建警报系统(避免警报疲劳)
- 参与值班轮班
- 使用构建智能警报
**提升系统可靠性 / 积极降低风险**
- 识别可靠性风险,帮助系统抵御故障
- 减少警报噪音 / 误报
- 增加可观测性覆盖范围(被监控服务的百分比)
- 提高 SLO 合规性
**仪器与遥测工程**
- 将可观测性嵌入应用程序
- 在代码中添加追踪/指标
- 标准化日志格式
查看英文原文
#### Company Description
Experian is a global data and technology company that drives opportunities for people and businesses around the world. We operate in diverse markets such as financial services, healthcare, automotive, agribusiness, insurance, and more. Experian invests in people and advanced new technologies to unlock the power of data. We have an incredible team of 25,200 employees in 32 countries.
Our uniqueness is valuing yours. Experian's people-centric, inclusive, and purpose-driven culture is recognized by numerous awards — including World’s Best Workplaces™ 2025 (Fortune's Top 25 global) and Great Place To Work™ in 26 countries, among others. Check out Experian Life on social media or explore our careers website to understand why. Experian is also proud to be an equal opportunity employer and an affirmative action employer. If you have a disability or need that requires accommodation, please let us know as soon as possible.
#### Job Description
A cloud observability engineer’s day is about making complex systems understandable, improving signal quality, and enabling faster, smarter debugging across teams.
- Check system heathy, review alerts/incidents.
- Triage alerts
- Investigate issues
- Improve observability instrumentation
- Build and improve dashboards
- Alert optimization
- Work with development teams and other engineering partners.
- Continuous improvements.
- Release support.
**Observability & Monitoring**
- Design and implement observability frameworks using metrics, logs, and distributed tracing
- Develop dashboards, alerts, and visualizations to monitor system health
- Standardize observability practices across engineering teams (logging, telemetry, tracing)
- Implement and manage native monitoring tools.
- Amazon CloudWatch (metrics, logs, alarms, dashboards)
- AWS X-Ray (distributed tracing)
- AWS Distro for OpenTelemetry (ADOT)
**Detection & Response**
- Build alerting systems (avoid alert fatigue)
- Participate in on-call rotations
- Build intelligent alerting using
**Improve System Reliability/ proactively reduce risk**
- Identify reliability risks to help harden systems against failure
- Reduction in alert noise / false positives
- Increased observability coverage (% of services instrumented)
- Improved SLO compliance
**Instrumentation & Telemetry Engineering**
- Embed observability into applications
- Add tracing/metrics into code
- Standardize logging formats
- Ensure all services are observable end-to-end
- Microservices (EKS, ECS)
- Serverless (Lambda, API Gateway)
- Data services (RDS, DynamoDB, S3)
**Collaboration Across Teams**
- Development
- Platform/infra teams
- Security & operations
#### Qualifications
**Requirements**
- SRE, DevOps, or Cloud Engineering
- Cloud platforms AWS
- Experience in AWS services including
- CloudWatch, X-Ray, Lambda
- ECS/EKS, API Gateway
- RDS, DynamoDB, S3
- Hands-on with experience with an observability tool
- Dynatrace
- Splunk
- Datadog
- OpenTelemetry
- Prometheus, Grafana
- AWS Distro for OpenTelemetry
- Strong understanding of:
- Containers & orchestration (Docker, Kubernetes)
- CI/CD pipelines
- Infrastructure as Code (Terraform, CloudFormation)
- Monitoring/observability tools
**Nice to have**
- Experience building observability platforms at scale in AWS
- Familiarity with multi-account AWS environments
- Experience with cost optimization for observability (logging/metrics ingestion)
- Experience in high-scale distributed system
#### Additional Information
At Serasa Experian, we believe that diversity is essential for a healthier and more innovative work environment, where everyone can share experiences and express their ideas. That’s why we promote several initiatives to support inclusive recruitment and the professional development of our people.
We also have our affinity groups, created to empower and support individuals from underrepresented groups: ExperianPride (LGBTQIAPN+ community), Ubuntu (racial equity), Women in Experian (gender equity), Aspire (people with disabilities), and Connecting Generations (generations).
Come be part of this transformation!
Experian Careers - Creating a better tomorrow together
[Find out what its like to work for Experian by clicking here](https://www.experian.com/careers/)