远程工作雷达

数据平台管理架构师

Data Platform Administration Architect

开发工程职能支持限定地区(需当地身份)
公司Wesco
薪资未公开
工作地点United States
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间昨天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

数据平台架构师负责企业数据平台的架构、管理、治理、运营卓越和持续改进。该职位负责云数据平台的稳定性、可扩展性、安全性、可用性和可支持性,主要专注于 Databricks、Azure 数据服务和企业平台运维。
理想的候选人需具备深厚的技术专长和强大的运营领导力,有丰富的生产环境管理、发布流程、变更治理、事件响应、服务管理及平台现代化项目经验。该人员将作为平台管理的技术负责人,与工程、基础设施、安全、治理和业务团队合作,确保数据平台的可靠和合规运作。

职责:

  • 领导企业数据平台的管理、架构、治理和运营支持。
  • 管理和支撑 Databricks、Azure Data Factory、Azure Data Lake Storage(ADLS)、Azure Synapse、Azure SQL 及相关 Azure 服务在生产及非生产环境中的运行。
  • 负责平台的可靠性、可用性、性能、可扩展性、安全性及运营卓越。
  • 制定并维护平台标准、环境策略、监控框架、操作流程和治理控制。
  • 监督变更管理与发布管理流程,包括变更请求(CR)、服务变更请求(SCR)、紧急变更请求(ECR)、部署计划、发布排期、回滚策略和发布后验证。
  • 参与变更咨询委员会(CAB)活动,提供技术评估、风险评估和部署准备建议。
  • 领导生产支持工作,包括处理事件(INC)、服务请求(RITM)、问题记录和重大事件管理流程。
  • 在关键生产事件中协调跨职能团队,确保服务及时恢复。
  • 执行根本原因分析(RCA)调查,并推动预防和纠正措施以减少重复性问题。
  • 编写和维护操作手册、知识文章、支持文档、升级流程和灾难恢复计划。
  • 建立并报告运营 KPI,包括平台正常运行时间、SLA 合规性、发布成功率、平均修复时间(MTTR)、事件趋势等。
查看英文原文

The Data Platform Administration Architect is responsible for the architecture, administration, governance, operational excellence, and continuous improvement of the enterprise data platform. This role owns the stability, scalability, security, availability, and supportability of cloud-based data platforms, with a primary focus on Databricks, Azure Data Services, and enterprise platform operations.
The ideal candidate combines deep technical expertise with strong operational leadership and has extensive experience managing production environments, release processes, change governance, incident response, service management, and platform modernization initiatives. This individual will serve as the technical lead for platform administration while partnering with engineering, infrastructure, security, governance, and business teams to ensure reliable and compliant data platform operations.
Responsibilities:

  • Lead the administration, architecture, governance, and operational support of enterprise data platforms.
  • Manage and support Databricks, Azure Data Factory, Azure Data Lake Storage (ADLS), Azure Synapse, Azure SQL, and related Azure services across production and non-production environments.
  • Own platform reliability, availability, performance, scalability, security, and operational excellence.
  • Define and maintain platform standards, environment strategies, monitoring frameworks, operational procedures, and governance controls.
  • Oversee Change Management and Release Management processes, including CRs, SCRs, ECRs, deployment planning, release scheduling, rollback strategies, and post-release validation.
  • Participate in CAB (Change Advisory Board) activities and provide technical assessment, risk evaluation, and deployment readiness recommendations.
  • Lead production support activities, including resolution of Incidents (INC), Service Requests (RITM), Problem Records, and Major Incident Management processes.
  • Coordinate cross-functional teams during critical production incidents and ensure timely restoration of services.
  • Perform Root Cause Analysis (RCA) investigations and drive preventative and corrective actions to reduce recurring issues.
  • Develop and maintain operational runbooks, knowledge articles, support documentation, escalation procedures, and disaster recovery plans.
  • Establish and report on operational KPIs including platform uptime, SLA compliance, release success rates, Mean Time to Resolution (MTTR), incident trends, and platform health metrics.
  • Implement and govern CI/CD processes, deployment automation, and infrastructure-as-code best practices.
  • Ensure compliance with enterprise security standards, access governance policies, audit requirements, and SOX controls.
  • Drive platform modernization initiatives, automation opportunities, cost optimization, and continuous service improvement efforts.
  • Mentor platform engineers and administrators while providing technical leadership and architectural direction.

Qualifications:

  • 10+ years of experience in enterprise data platforms, platform engineering, cloud infrastructure, platform administration, or data operations.
  • Hands-on experience administering and supporting Databricks and Azure data platforms in enterprise production environments.
  • Strong expertise with:
  • Databricks
  • Azure Data Lake Storage (ADLS)
  • Azure Synapse Analytics
  • Azure SQL
  • Azure Data Factory (ADF)
  • Azure DevOps
  • Azure IAM (RBAC, Service Principals, Managed Identities)
  • Azure Monitor and Log Analytics
  • Proven experience as a Data Platform Administrator, Databricks Administrator, Azure Platform Administrator, Platform Architect, or Platform Operations Lead, not solely a data engineer or application developer.
  • Deep understanding of standing up, configuring, securing, governing, and supporting production and non-production Databricks environments.
  • Experience managing platform provisioning, environment strategy, capacity planning, workload management, performance tuning, cost optimization, and operational support.
  • Strong experience implementing and managing platform monitoring, observability, alerting, and operational dashboards.
  • Proven experience managing enterprise production support processes, including:
  • Incidents (INC)
  • Service Requests (RITM)
  • Change Requests (CR)
  • Standard Changes (SCR)
  • Emergency Changes (ECR)
  • Problem Management
  • Hands-on experience coordinating and leading enterprise release management activities, including release planning, deployment execution, rollback planning, deployment validation, and post-release support.
  • Strong knowledge of ITIL-aligned operational processes, including:
  • Incident Management
  • Change Management
  • Release Management
  • Problem Management
  • Service Request Management
  • Knowledge Management
  • Major Incident Management
  • Experience resolving and leading response efforts for critical production incidents (P1/P2/P3), including executive communications and stakeholder coordination.
  • Demonstrated experience creating and presenting comprehensive Root Cause Analysis (RCA) reports and corrective action plans following production incidents.
  • Experience establishing support operating models, on-call rotations, escalation procedures, and operational readiness reviews.
  • Strong experience with ServiceNow or equivalent ITSM platforms for managing incidents, changes, requests, approvals, releases, and audit requirements.
  • Experience developing operational metrics, service-level reporting, trend analysis, and platform health dashboards.
  • Experience designing scalable platform controls and operational processes to support enterprise data volumes and growth.
  • Strong experience with CI/CD automation, deployment pipelines, Infrastructure as Code (IaC), and Azure DevOps best practices.
  • Experience with cloud governance, platform security, access management, and compliance controls.
  • Experience supporting SOX-compliant, regulated, or audit-driven environments.
  • Ability to partner effectively with infrastructure, security, networking, application, and business teams to drive platform reliability and operational excellence.
  • Proven leadership experience managing platform, operations, or engineering teams.
  • Strong communication, stakeholder management, leadership, and problem-solving skills.
  • Bachelor's degree in Computer Science, Information Systems, Engineering, or related field.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位