现场运营总监
Director, Live Operations
总监,实时运营
远程职位
关于Reveleer
Reveleer 提供一个统一的平台,涵盖风险调整、质量改进、临床智能和会员管理,为医疗计划和提供者组织在价值医疗的复杂环境中提供支持。全国80多家客户组织信赖该平台,该平台将数据、分析和智能工作流自动化整合到一个受控系统中,旨在支持诊断、质量指标和提交的可追溯文档。凭借监管专业知识和透明的人机协同AI核心,Reveleer 支持致力于提升护理质量、加强文档完整性并保持运营准备以自信应对审计的组织。
职位概述
实时运营总监负责技术生产运维职能,确保CDM管理的客户数据服务的可用性、可靠性、可支持性和持续运行。该职位负责生产数据工作流的24x7运营模式,并对重大运营事件、服务恢复、问题管理和持续可靠性改进负主要责任。
该角色领导一个涵盖云生产支持、数据管道运维、可观测性、基础设施自动化、工作流恢复、安全数据传输、访问/连接性、运维工具和自动化的运维工程团队。这是一个需要亲力亲为的技术领导职位:总监必须能够理解并挑战AWS、Terraform/基础设施即代码、数据管道、API、作业编排、监控和生产自动化中的工程方法。
主要职责
24x7生产可靠性与服务责任
- 负责CDM管理的生产数据服务和工作流的24x7运营健康和可靠性。
- 作为严重程度1和严重程度2生产事件的负责人和升级责任人。
- 制定SLA/SLO、升级路径、值班/覆盖模型、服务健康度指标和运营绩效期望。
- 推动服务恢复、沟通协调和纠正措施的跟进。
事件与问题管理
- 负责CDM事件和问题管理专业领域,包括严重程度定义、事件指挥、升级、根本原因分析、事件后回顾和纠正措施。
- 跟踪重复性故障并使用MTTA、MTTR、可用性等指标进行分析。
查看英文原文
Director, Live Operations
Remote Opportunity
About Reveleer
Reveleer delivers a unified platform spanning risk adjustment, quality improvement, clinical intelligence, and member management for health plans and provider organizations navigating the complexity of value-based care. Trusted by 80+ customer organizations nationwide, the platform integrates data, analytics, and intelligent workflow automation into one governed system designed to support traceable documentation across diagnoses, quality measures, and submissions. With regulatory expertise and transparent, human-in-the-loop AI at its core, Reveleer supports organizations working to advance care quality, strengthen documentation integrity, and sustain the operational readiness needed to navigate audits with confidence.
Position Summary
The Director, Live Operations, leads the technical production-operations function responsible for the availability, reliability, supportability, and continuous operation of CDM-managed customer data services. This leader owns the 24x7 operating model for production data workflows and is the accountable leader for major operational incidents, service restoration, problem management, and continuous reliability improvement.
The role leads an operations-engineering organization spanning cloud production support, data-pipeline operations, observability, infrastructure automation, workflow recovery, secure data movement, access/connectivity, operational tooling, and automation. This is a hands-on technical leadership role: the Director must be able to understand and challenge engineering approaches across AWS, Terraform/infrastructure-as-code, data pipelines, APIs, job orchestration, monitoring, and production automation.
Key Responsibilities
24x7 Production Reliability & Service Ownership
- Own the 24x7 operational health and reliability of CDM-managed production data services and workflows.
- Serve as accountable leader and escalation owner for Sev-1 and Sev-2 production incidents.
- Define SLAs/SLOs, escalation paths, on-call/coverage models, service-health measures, and operational performance expectations.
- Drive service restoration, communication coordination, and corrective-action follow-through.
Incident & Problem Management
- Own CDM incident and problem-management disciplines, including severity definitions, incident command, escalation, RCA, post-incident review, and corrective actions.
- Track recurring failures and use MTTA, MTTR, availability, incident volume, and recurrence metrics to drive systemic improvement.
Cloud & Infrastructure Operations
- Provide technical leadership for production services operating in AWS and related enterprise environments.
- Partner with Engineering on infrastructure-as-code using Terraform or comparable tooling, including repeatable configuration, deployment, and recovery.
- Guide operational issues involving IAM, networking/connectivity, secure file transfer, storage, compute, logging, monitoring, and cloud dependencies.
DevOps, Automation & Operations Engineering
- Lead an automation-first strategy to eliminate repetitive manual work, fragile handoffs, and key-person dependencies.
- Drive scripting, orchestration, automated validation, job recovery, exception handling, and self-healing patterns where appropriate.
- Partner with Data Engineering on CI/CD, APIs, ETL/data pipelines, file movement, deployment/support patterns, and production automation.
- Apply AI-assisted monitoring, troubleshooting, documentation, and workflow automation where appropriate.
Observability & Operational Readiness
- Establish monitoring, logging, alerting, and operational dashboards that provide actionable visibility into production health.
- Define production-readiness gates for workflows transitioning from implementation or engineering into Live Operations.
- Require current runbooks, SOPs, recovery procedures, escalation paths, dependency maps, ownership, and cross-trained coverage before production handoff.
Cross-Functional & Team Leadership
- Define clear operating boundaries among Live Operations, Data Engineering, Data Management, Clinical Intelligence, Product, and IT.
- Coordinate technical response across teams during production incidents and complex operational issues.
- Build and lead a geographically distributed technical operations team with a culture of urgency, transparency, documentation, collaboration, and measurable improvement.
Required Qualifications
- 8+ years in production operations, cloud/platform operations, DevOps, SRE, operations engineering, data platform operations, or a comparable high-availability technical environment.
- 4+ years leading technical production operations, DevOps, SRE, platform-support, or operations-engineering teams.
- Demonstrated accountability for business-critical production systems operating under extended-hours or 24x7 support models.
- Demonstrated leadership of Sev-1/Sev-2 or equivalent major incidents, including incident command, restoration, RCA, problem management, and corrective-action follow-through.
- Strong working knowledge of AWS production environments, including IAM, networking/connectivity, logging/monitoring, cloud dependencies, and operational troubleshooting.
- Demonstrated experience with Terraform or comparable infrastructure-as-code technologies and repeatable infrastructure deployment/recovery practices.
- Strong technical understanding of APIs, ETL/data pipelines, secure file transfer, workflow/job orchestration, SQL, automation/scripting, and production integration patterns.
- Demonstrated experience establishing observability, monitoring, alerting, runbooks, operational dashboards, and measurable reliability practices.
- Demonstrated success automating manual production processes and reducing key-person dependencies through tooling, scripting, orchestration, or platform improvements.
- Experience managing availability, SLA/SLO attainment, MTTA, MTTR, incident volume, recurring failures, and automation coverage.
- Experience leading geographically distributed technical teams and effective on-call, escalation, and coverage models.
- Ability to lead across Engineering, Product, IT, implementation, and customer-facing organizations during production incidents and reliability initiatives.
Preferred Qualifications
- Experience operating high-volume healthcare, financial-services, SaaS, or other regulated production data environments.
- Experience with MuleSoft, Flatfile, or comparable enterprise integration and data-ingestion platforms.
- Experience with CI/CD tooling, cloud observability platforms, workflow orchestration, and automated recovery patterns.
- Experience applying AI-assisted tooling to production support, incident analysis, documentation, or operational automation.
- Healthcare payer data, CMS submissions, EDI/X12, enrollment, claims, risk-adjustment, or related domain knowledge.
Leadership Profile
- Technically credible with engineers and comfortable going deep on architecture and production failure modes.
- Calm, structured, and decisive during high-severity incidents.
- Automation-first and focused on eliminating avoidable manual/key-person dependencies.
- Transparent in operational communication and documentation.
- Collaborative across organizational boundaries while maintaining clear accountability.
- Metrics-driven and focused on systemic reliability improvement rather than repeated tactical recovery.
WHAT YOU’LL RECEIVE:
• Competitive salary
• Medical, Dental and Vision benefits
• 401k match
• Generous PTO plan
Our compensation reflects the cost of labor across several US geographic markets. Pay is based on several factors including market location and may vary depending on job-related knowledge, skills, and experience.
Reveleer E-Verifies all new hires.
Reveleer is an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender identity, sexual orientation, age, marital status, veteran status, disability status or genetic information, in compliance with applicable federal, state and local law.
Originally posted on Himalayas