远程工作雷达

高级数据工程师

Sr Data Engineer

开发工程未标注地域
公司Protective
薪资未公开
工作地点Birmingham, AL
地域资格未标注地域
时区要求无特别要求
用工类型Full Time
发布时间未知
数据来源Lever
前往企业招聘页投递 →

我们的工作影响数百万人的生活,你可以成为其中一员。
我们帮助客户抵御生活中的不确定性。无论你在公司哪个部门工作,你都将帮助客户在他们最需要的时候提供保障和安心。

Protective 正在寻找数据工程师,在 Voyager 上设计、构建和运营生产数据流水线,Voyager 是我们基于 Azure 的 Databricks 数据湖仓。你将负责通过分层架构(medallion architecture)的数据全流程——将源数据导入青铜层,进行清洗、验证和合规处理,然后在黄金层发布可消费的可信数据产品。

这个职位位于工程、分析和平台运维的交汇点。你将编写生产级别的 Python 和 SQL,执行数据契约,并帮助 Protective 将数据视为具有指定负责人和真实用户的的产品。你将被分配到一个交付小组,该小组对数据产品从头到尾负责,而不是从队列中处理工单。

在 Voyager 上,分层架构的名称为 Raw、Prep 和 Prod。它们与青铜、白银和黄金相对应,在本描述中可互换使用。

员工福利:
我们致力于通过广泛的福利计划保护员工及其家庭的健康。除了提供全面的医疗、牙科和视力保险外,我们还通过心理健康福利和员工援助计划支持员工的情绪健康。工作与生活的平衡很重要,Protective 提供多种带薪休假福利(例如带薪年假、带薪育儿假、短期残疾假期和文化纪念日)。员工的财务健康与身体和情绪健康同样重要。一些财务健康福利包括医疗账户的贡献、养老金计划和有公司匹配的 401(k) 计划。所有员工都被鼓励通过参与 ProHealth Rewards——Protective 的提升健康状况并赚取现金奖励的平台——来保护自身的整体健康。

某些福利的资格可能根据职位不同而有所差异,具体以公司福利计划条款为准。

残疾人申请者的便利措施:
如果你因残疾需要便利措施以完成申请和招聘流程,请发送邮件至 eric.hess@protective.com。此信息将被保密,并仅用于确定适当的便利措施。

查看英文原文

The work we do has an impact on millions of lives, and you can be a part of it.
We help protect our customers against life’s uncertainties. Regardless of where you work within the company, you’ll be helping provide protection and peace of mind when our customers need it most.

Protective is looking for Data Engineers to design, build, and operate production data pipelines on Voyager, our Databricks lakehouse on Azure. You will own the end-to-end flow of data through a medallion architecture — ingesting source data into the Bronze layer, applying cleansing, validation, and conformance in Silver, and publishing trusted, consumption-ready data products in Gold.

This role sits at the intersection of engineering, analytics, and platform operations. You will write production-grade Python and SQL, enforce data contracts, and help Protective treat data as a product with named owners and real consumers. You will be embedded on a delivery pod that owns its data products end to end, rather than servicing tickets from a queue.

On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description.
Employee Benefits:  
We aim to protect the wellbeing of our employees and their families with a broad benefits offering. In addition to offering comprehensive health, dental and vision insurance, we support emotional wellbeing through mental health benefits and an employee assistance program. Work/life balance is important and Protective offers a variety of paid time away benefits (e.g., paid time off, paid parental leave, short-term disability, and a cultural observance day). The financial health of our employees is just as important as physical and emotional health.  Some of the financial wellbeing benefits include contributions to healthcare accounts, a pension plan, and a 401(k) plan with Company matching. All employees are encouraged to protect their overall wellbeing by engaging in ProHealth Rewards, Protective’s platform to improve wellbeing while earning cash rewards.

Eligibility for certain benefits may vary by position in accordance with the terms of the Company’s benefit plans.

Accommodations for Applicants with a Disability:
If you require an accommodation to complete the application and recruitment process due to a disability, please email eric.hess@protective.com. This information will be held in confidence and used only to determine an appropriate accommodation for the application and recruitment process.

Please note that the above email is solely for individuals with disabilities requesting an accommodation.  General employment questions should not be sent through this process.

We are proud to be an equal opportunity employer committed to being inclusive and attracting, retaining, and growing an inclusive workforce.

Key Responsibilities
•    Design, develop, and maintain production data pipelines on Databricks using Python, SQL, Apache Spark, and Delta Lake.
•    Build Bronze-layer ingestion that reliably captures data from APIs, relational databases, flat files, cloud storage, and SaaS platforms — using dlt (dltHub) and Databricks-native ingestion where each fits — including incremental loading, pagination, watermarking, state management, and replay after failure.
•    Develop Silver-layer transformations in dbt and Python over Delta Lake that cleanse, standardize, type, deduplicate, validate, conform, and enrich data so that it is reusable across domains. A meaningful share of this role is making messy source data trustworthy.
•    Create Gold-layer data products: dimensional models, slowly changing dimensions, fact and bridge tables, aggregates, and serving tables aligned to how consumers actually query.
•    Produce and maintain the curated datasets ML engineering trains and serves models from — feature and training tables that are versioned and reproducible, not one-off extracts.
•    Author and maintain data contracts using the Open Data Contract Standard (ODCS) — schema with real semantics, named owner, known consumers, quality rules, and freshness expectations — and assess backward compatibility before every change.
•    Implement data quality as code: uniqueness and not-null on keys at minimum, plus referential, accepted-value, freshness, and custom business-rule tests, surfaced to producers and consumers rather than buried in logs.
•    Orchestrate ingestion and transformation as assets in Dagster, deployed to Dagster Cloud, and operate what you build across development, branch, and production deployments — schedules and sensors, asset dependencies, backfills, and run observability.
•    Apply governance through Unity Catalog — catalogs, schemas, external locations, grants, row- and column-level security, and lineage — and handle credentials through Azure Key Vault rather than in code.
•    Implement incremental and merge-based processing with Delta Lake (MERGE, schema evolution, time travel, OPTIMIZE) and tune Spark jobs, table layouts, and compute for performance and cost.
•    Troubleshoot production failures, data-quality issues, source-system changes, and late-arriving or duplicate data — including backfills and recovery — and take part in the pod’s on-call rotation for the pipelines it owns, with root-cause analysis that closes the gap rather than reopening the ticket.
•    Build and maintain CI/CD for data assets in Azure DevOps — automated tests and CI checks on dlt, dbt, and Dagster changes, promotion from development through branch deployments to production, and releases that are repeatable and auditable.
•    Instrument what you own for observability: freshness, volume, quality, latency, and cost, with alerting tied to the SLAs and SLOs your contract commits to instead of depending on someone noticing.
•    Work inside the platform’s control expectations — least-privilege access, secrets in Azure Key Vault, change management through pull request and pipeline, and audit evidence that falls out of the deployment path rather than being reconstructed later.
•    Participate in code review and document architecture, runbooks, and data products so others can discover, trust, and reuse them.
•    Work with data architects, analysts, product owners, and business stakeholders to translate requirements into maintainable data solutions.

Qualifications
Required Qualifications

•    Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered.
•    3+ years building and supporting production data pipelines in a cloud data platform environment.
•    Strong hands-on Python and SQL. Both are used daily and neither substitutes for the other.
•    Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake tables, MERGE, and incremental load patterns.
•    Practical understanding of medallion / multi-layer lakehouse design, and the judgment to say what belongs in Bronze versus Silver versus Gold.
•    Experience ingesting data from APIs, relational databases, files, or SaaS applications, including the incremental and state-management problems that come with it.
•    Working knowledge of dimensional modeling — grain, keys, facts and dimensions, slowly changing dimensions — and of ELT design patterns and data quality practice.
•    Experience with orchestration and scheduling using Dagster, Databricks Workflows, Airflow, Azure Data Factory, or similar.
•    Git-based source control, pull request review, automated testing, and CI/CD as normal practice — Azure DevOps or comparable.
•    Experience troubleshooting production data failures, performance bottlenecks, and source-system changes.
•    Experience with pipeline monitoring and alerting, and a working understanding of what a freshness or quality SLA means once real consumers depend on it.
•    Ability to explain technical designs and trade-offs to both technical and non-technical partners.

Preferred Qualifications

•    Databricks certification (Data Engineer Associate or Professional) or equivalent demonstrated depth.
•    Unity Catalog experience: catalogs, schemas, volumes, external locations, storage credentials, permissions, and lineage.
•    dbt on Databricks, or another transformation framework used alongside Spark.
•    Python-based modeling frameworks over Delta Lake, and experience implementing Type 2 history, surrogate keys, and merge strategies in code.
•    Experience with a declarative Python ingestion framework such as dlt (dltHub), Airbyte, Meltano, or Fivetran.
•    Dagster experience specifically, including assets, asset checks, sensors, schedules, and branch deployments.

本页面信息整理自 Lever,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位