FBS AIOps 工程师
FBS AIOps Engineer
FBS – 农民业务服务是农民业务的一部分,旨在建立一个全球化的策略,用于识别、招聘、雇佣和留住顶尖人才。通过结合国际影响力与美国的专业知识,我们打造多元化且高效能的团队,使他们在当今竞争激烈的市场中能够蓬勃发展。
我们认为,每个成功企业的基础在于拥有具备正确技能的正确人才。这就是我们的作用——帮助农民打造一支能够持续稳定取得成果的优秀团队。
由于我们没有当地的法律实体,我们与Capgemini合作,由其作为雇主代表。Capgemini负责管理当地的薪资和福利。
要求
简介:
AIOps工程师是集中式AIOps平台团队的一员,负责设计、构建和运营一个共享的AIOps技术栈,以支持企业内部的IT运维、SRE和基础设施团队。该角色与下游运维团队紧密合作,提供可操作的警报、智能事件关联和运行手册自动化,同时确保AIOps平台具备可扩展性、可靠性,并符合企业标准。AIOps工程师不直接嵌入单一运维团队,而是通过可重用模式、平台功能和AI驱动的洞察力来支持多个团队。
职责:
- 设计、实施并维护一个集中式AIOps平台,从多个来源(指标、日志、追踪、事件、工单)获取可观测性和操作数据。
- 构建并维护与监控、ITSM、CI/CD和云平台的集成。
- 确保平台具备可扩展性、弹性和性能,以支持全企业范围的应用。
- 维护标准化的数据模型、实体定义和服务映射,以支持事件关联和根本原因分析。
- 可操作警报与事件关联
- 与运维团队合作,定义以可操作性和业务影响为导向的警报需求,而非信号数量。
- 实现并优化事件关联功能,减少警报噪音,并在基础设施和应用层之间对相关事件进行分组。
- 利用拓扑结构、依赖关系映射和上下文增强,生成与服务和客户体验对齐的有意义警报。
- 持续评估警报效果,并根据运维反馈调整关联逻辑。
- 启用运行手册自动化
查看英文原文
FBS – Farmer Business Services is part of Farmers operations with the purpose of building a global approach to identifying, recruiting, hiring, and retaining top talent. By combining international reach with US expertise, we build diverse and high-performing teams that are equipped to thrive in today’s competitive marketplace.
We believe that the foundation of every successful business lies in having the right people with the right skills. That is where we come in—helping Farmers build a winning team that delivers consistent and sustainable results.
Since we don’t have a local legal entity, we’ve partnered with Capgemini, which acts as the Employer of Record. Capgemini is responsible for managing local payroll and benefits.
Requirements
Summary:
The AIOps Engineer is a member of the centralized AIOps platform team responsible for designing, building, and operating a shared AIOps technology stack that supports IT Operations, SRE, and Infrastructure teams across the enterprise. This role partners closely with downstream operations teams to deliver actionable alerts, intelligent event correlation, and runbook automation, while ensuring the AIOps platform is scalable, reliable, and aligned to enterprise standards. Rather than embedding directly within a single operations team, the AIOps Engineer enables multiple teams through reusable patterns, platform capabilities, and AI driven insights.
Responsibilities:
- Design, implement, and maintain a centralized AIOps platform that ingests observability and operational data from multiple sources (metrics, logs, traces, events, tickets). • Build and maintain integrations with monitoring, ITSM, CI/CD, and cloud platforms.
- Ensure platform scalability, resiliency, and performance for enterprise wide adoption.
- Maintain standardized data models, entity definitions, and service mappings to support correlation and root cause analysis.
- Actionable Alerting & Event Correlation • Partner with operations teams to define alerting requirements that focus on actionability and business impact, not signal volume.
- Implement and tune event correlation capabilities to reduce alert noise and group related events across infrastructure and application layers.
- Leverage topology, dependency mapping, and contextual enrichment to produce meaningful alerts aligned to services and customer experience.
- Continuously evaluate alert effectiveness and adjust correlation logic based on operational feedback.
- Enable runbook automation workflows that support incident response, investigation, and recovery.
- Integrate alerts and correlated events with automated actions, enrichment steps, and remediation activities (with appropriate safeguards). • Support standardized incident workflows by integrating AIOps outcomes into ITSM tools and collaboration platforms.
- Help operational teams transition manual responses into repeatable, automated runbooks over time.
- Act as a consultative partner to operations, SRE, and platform teams. • Gather use cases, translate operational needs into AIOps solutions, and prioritize work within the centralized backlog.
- Enable self service capabilities while maintaining governance and guardrails. • Provide guidance, documentation, and best practices to ensure consistent adoption across teams.
- Monitor accuracy and effectiveness of AIOps features such as anomaly detection, correlation, and RCA.
- Incorporate feedback loops from incident outcomes and operator input. • Ensure solutions align to enterprise standards for security, compliance, and AI governance.
- Contribute to platform roadmaps, technical standards, and long term AIOps strategy
Skills & Tools:
- Dynatrace - Intermediate (2-4 Years) MUST
- Copilot Studio - Entry Level (1-2 Years) MUST
- ITIL Foundations
- Service Now – Desirable
- Snowflake – Desirable
- Python – Entry Level - Desirable
Benefits
This position comes with a competitive compensation and benefits package.
· A competitive salary and performance-based bonuses.
· Comprehensive benefits package.
· Flexible work arrangements (remote and/or office-based).
· You will also enjoy a dynamic and inclusive work culture within a globally renowned group.
· Private Health Insurance.
· Paid Time Off.
· Training & Development opportunities in partnership with renowned companies.
Originally posted on Himalayas