高级数据工程师
Senior Data Engineer
我们是高精度真实世界证据(RWE)生成领域的市场领导者。Verantos 的 RWE 平台整合异构的真实世界数据源,并生成符合监管和报销要求的准确证据。该平台利用数据科学和人工智能,以及电子健康记录(EHR)等先进数据源,生成能够支持复杂临床研究的 RWE。全球一些最大的生物制药公司都是 Verantos 的客户。
我们使用以 AWS 为中心的异构技术栈,涵盖数据处理、AI、工作流和分析。我们的团队是跨职能且协作的,汇聚了工程、产品、设计、QA 和临床领域专家,以实现有意义的真实世界影响。
职位描述
推动 Verantos 证据平台的数据来自真实世界的临床系统——混乱、不一致且不断变化。我们需要一位高级数据工程师,懂得如何构建能优雅处理这种混乱的管道,而不是每次遇到意外情况都去灭火。
这是负责每季度交付数据产品的团队中的高级职位。你将确定数据摄取、转换和质量检查的技术方向,注重可自我运行的系统。同样重要的是能够超越管道本身思考:优秀的候选人理解数据对依赖它的研究人员意味着什么,并将这种视角带入他们所做的工程决策中。
这是一个完全远程、面向美国的职位。
职责
- 领导数据平台架构的设计与演进,建立团队可依托的模式和标准。
- 构建并运维生产级数据管道,可靠且大规模地摄取和转换高差异的真实世界临床数据。
- 从一开始就设计自动化:能够检测问题、优雅恢复,并在无需人工干预的情况下运行的管道。
- 参与季度数据产品发布,与产品、临床和客户成功团队紧密合作以达成承诺。
- 构建反映下游用户不断变化需求的数据质量测试。
- 通过代码审查、架构决策和共享标准,指导和提升其他数据工程师。
- 积极使用并倡导 AI 工具,以提高团队的开发速度和代码质量。
查看英文原文
Overview
Verantos is the market leader in high-accuracy real-world evidence (RWE) generation. The Verantos RWE platform integrates heterogeneous real-world data sources and generates evidence with the accuracy necessary for regulatory and reimbursement use. The platform leverages data science and artificial intelligence, along with advanced data sources such as electronic health records (EHR), to generate RWE capable of supporting complex clinical studies. Some of the largest biopharma companies in the world are Verantos customers.
We use a heterogeneous AWS-centric tech stack with data processing, AI, workflows, and analytics. Our teams are cross-functional and collaborative, bringing together engineering, product, design, QA, and clinical domain experts to deliver meaningful, real-world impact.
Job Description
The data that powers Verantos's evidence platform comes from real-world clinical systems — messy, inconsistent, and constantly changing. We need a Senior Data Engineer who knows how to build pipelines that handle that chaos gracefully, not one who fights fires every time something unexpected arrives.
This is a senior role on the team responsible for shipping our data product every quarter. You will set the technical direction for how we ingest, transform, and quality-check data at scale, with an eye toward systems that run themselves. Just as important is the ability to think beyond the pipeline: the best candidate understands what the data means to the researchers who depend on it, and brings that perspective into the engineering decisions they make.
This is a fully remote, US-based role.
Responsibilities
- Lead the design and evolution of the data platform architecture, establishing patterns and standards the team builds on.
- Build and operate production-grade data pipelines that ingest and transform high-variance, real-world clinical data reliably and at scale.
- Design for automation from the start: pipelines that detect problems, recover gracefully, and surface issues without requiring manual intervention to run.
- Contribute to quarterly data product releases, working closely with product, clinical, customer success teams to meet commitments.
- Build data quality tests that reflect the evolving needs of our downstream consumers.
- Mentor and elevate other data engineers through code review, architecture decisions, and shared standards.
- Actively use and advocate for AI tools that improve the team's development velocity and code quality.
Qualifications
- 8+ years in data engineering, with experience at a technical lead level.
- Production experience with Snowflake and dbt as primary data platform tools.
- Strong Python skills for building and maintaining data pipelines.
- Has built resilient pipelines on irregular, high-variance data sources and knows what it takes to keep them running without babysitting.
- Thinks in systems: designs for observability, failure recovery, and automation.
- Can engage meaningfully with the business and domain context around the data, not just the engineering.
- Uses AI tools actively in their own work and is curious about applying them within the pipeline, particularly for data quality monitoring and anomaly detection at scale.
- Communicates clearly and works well across engineering, product, and clinical stakeholders.
Nice to Have
- Familiarity with OMOP CDM — not required, but it matters here more than most places.
- Experience with EHR data or other clinical datasets.
- Familiarity with other healthcare data standards such as HL7 or FHIR.
- Experience with data observability tooling in production environments.
Compensation
The base salary range for this position is $150,000–$220,000, depending on experience.