远程工作雷达

高级可观测性工程师 II

Sr. Observability Engineer II

开发工程限定地区(需当地身份)
公司NextGen Healthcare
薪资未公开
工作地点India
地域资格限定地区(需当地身份)
时区要求日间重叠约 6 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 India 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

职位描述:
高级可观测性工程师 II 是一个专业工程角色,负责 Dynatrace 可观测性平台的专家级支持。该职位的主要职责是通过专注于信号质量工程和操作噪声减少,提升公司主动监控能力。该职位在 Hosting Operations、站点可靠性工程(SRE)和其他工程团队之间担任关键的技术联络人,以确保我们面向客户的医疗 SaaS 环境的稳定性和可靠性。
· 设计和维护 Dynatrace 多租户环境配置,包括管理区域、告警配置文件、指标和自定义事件、标签策略以及访问边界。

  • 负责端到端的告警生命周期,定义所有权、命名标准、严重性映射、告警处置分类法和生产就绪标准。
  • 通过高级调优、阈值、相关性、抑制和告警退役,系统性地消除顶级问题模式中的告警疲劳,同时保持对真实客户影响问题的全面覆盖。
  • 设计并实现合成监控和以业务为中心的服务健康模型,用于关键客户旅程和访问路径,确保在服务和业务术语中检测退化情况,而不是原始基础设施指标。
  • 编写并标准化 Dynatrace 查询语言(DQL)查询、笔记本和仪表板,用于运营、管理层和可靠性报告,为团队建立最佳实践开发标准。
  • 与工具管理员合作,使用配置即代码(configuration-as-code)实践来管理可观测性配置,确保监控设置版本化、经过同行评审并可重复。
  • 与工具、自动化和 SRE 团队合作,增强告警到工单的丰富信息(如 Salesforce ITSM 集成),提供上下文和可能原因,减少初步排查时间并实现自动化修复。
  • 与 SQL、平台、操作系统和 API 专家合作,解决监控盲点,确保领域特定告警有意义,并验证告警有坚实的运行手册支持。
  • 支持服务级别(P1–P4)度量,通过建立企业问责制的报告基线,确保降噪变更与严格的运维保障措施相结合。
  • 维护全面的参考材料、标准操作流程和可观测性标准文档。
查看英文原文

Job Description:
The Sr. Observability Engineer II is a specialized engineering role serving as the subject matter expert for the Dynatrace observability platform. The primary purpose of this role is to mature the firm’s proactive monitoring capabilities by focusing on signal quality engineering and operational noise reduction. This position functions as a key technical liaison between Hosting Operations, Site Reliability Engineering (SRE), and other engineering teams to ensure the stability and reliability of our client-facing healthcare SaaS environment.
· Engineer and maintain the Dynatrace multi-tenant environment configuration, including management zones, alerting profiles, metric and custom events, tagging strategies, and access boundaries.

  • Own the end-to-end alert lifecycle, defining ownership, naming standards, severity mapping, alert disposition taxonomies, and production-readiness criteria.
  • Systematically eliminate alert fatigue across top problem patterns utilizing advanced tuning, thresholds, correlation, suppression, and alert retirement, while maintaining robust coverage for real, client-impacting issues.
  • Design and implement synthetic monitors and business-centric service-health models for critical client journeys and access paths, ensuring degradation is detected in service and business terms rather than raw infrastructure metrics.
  • Author and standardize Dynatrace Query Language (DQL) queries, notebooks, and dashboards for operations, leadership, and reliability reporting, establishing best-practice development standards for the team.
  • Collaborate with tools administrators to manage observability configurations using configuration-as-code practices, ensuring monitoring setups are versioned, peer-reviewed, and reproducible.
  • Partner with Tools, Automation and SRE teams to enhance alert-to-ticket enrichment (such as Salesforce ITSM integrations) with context and probable cause, minimizing triage times and enabling automated remediation.
  • Partner with SQL, platform, OS, and API specialists to resolve monitoring blind spots, ensure domain-specific alerts are meaningful, and verify that alerts are backed by robust runbooks.
  • Support service-level (P1–P4) measurements by establishing reporting baselines for corporate accountability, ensuring noise-reduction changes are paired with strict operational guardrails.
  • Maintain comprehensive reference materials, standard operating procedures, and observability standards documentation in Confluence to support continuous team enablement and operational onboarding.
  • Perform other duties that support the overall objective of the position.

Education Required:

  • Bachelor’s Degree in Computer Science, Information Technology, or a related technical field.
  • Or, any combination of education and experience which would provide the required qualifications for the position.

Experience Required:

  • 7+ years of experience in observability, application performance monitoring (APM), monitoring engineering, site reliability engineering (SRE), or a closely related technology infrastructure field.
  • Demonstrable hands-on expertise with the Dynatrace platform, including Davis AI, DQL, dashboards, notebooks, management zones, alerting profiles, metric/custom events, and synthetic monitors.
  • Proven track record of designing, configuring, and tuning alerting systems at enterprise scale to successfully reduce noise while maintaining robust detection capabilities.
  • Experience integrating monitoring systems with enterprise ITSM/ticketing systems (such as Salesforce) for automated ticket routing and enrichment.
  • Experience working within highly regulated hosting environments (e.g., healthcare, HIPAA/HITRUST, SOC 2, or ISO 27001).

License/Certification Required:
· Dynatrace certification (or commitment to obtain certification within 1 year of hire).
Knowledge, Skills & Abilities:

  • Knowledge of: Strong working knowledge of AWS infrastructure, Windows and Linux operating systems, and SQL Server database concepts. Advanced understanding of cloud-native observability frameworks, APM tools (specifically Dynatrace), and alert lifecycle management. Strong command of cloud computing (AWS), systems integration, database operations, and scripting (e.g., Python or PowerShell) to support automation.
  • Skill in: Excellent problem-solving capabilities. Strong written and verbal communication skills. Highly organized, detail-oriented, and self-driven, with a strong ownership mindset toward maintaining signal quality, reducing operational toil, and defending client service-level agreements.
  • Ability to: Proven ability to translate complex technical infrastructure metrics into direct service and business impacts. Ability to collaborate effectively across operations command analysts, SRE partners, database specialists, and leadership.

The company has reviewed this job description to ensure that essential functions and basic duties have been included. It is intended to provide guidelines for job expectations and the employee's ability to perform the position described. It is not intended to be construed as an exhaustive list of all functions, responsibilities, skills and abilities. Additional functions and requirements may be assigned by supervisors as deemed appropriate. This document does not represent a contract of employment, and the company reserves the right to change this job description and/or assign tasks for the employee to perform, as the company may deem appropriate.
NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

客户交付经理

NextGen HealthcareGabon, GeorgiaFull Time今天
职能支持未标注地域

首席数据库工程师

NextGen HealthcareGabonFull Time2 天前
开发工程未标注地域日间重叠约 2 小时,需偶尔早起或晚睡

← 返回全部职位