AWS数据工程师
AWS Data Engineer
ABOUT THE ROLE
这是医疗保健行业数据工程团队中一个完全动手的个人贡献者角色。你将负责数据从主框架源流通过AWS摄入到提供干净、消费者就绪数据集的全过程。你的工作直接支持大规模的下游分析和运营系统。
WHAT YOU'LL DO
- 使用现代数据管道工具将主框架源数据流处理并导入AWS S3。
- 使用AWS Glue设计和实现ETL/ELT工作流,对数据进行整理和转换。
- 对摄入的数据集执行数据对账、验证和质量检查。
- 通过AWS Aurora和RDS PostgreSQL数据库提供干净、消费者就绪的数据集。
- 管理S3存储,包括数据保留策略、归档策略和生命周期管理。
- 使用GitHub Actions和CI/CD流水线自动化和部署工作流。
- 利用AI工具提高工程效率并自动化数据工作流。
WHAT WE'RE LOOKING FOR
- 5年以上专业数据工程经验,交付过数据管道、ETL/ELT工作流或数据平台解决方案。
- 具有在生产环境中使用AWS Glue进行ETL/ELT设计和实现的实际经验。
- 熟悉AWS S3、RDS和Aurora在生产环境中的使用。
- 有使用Kafka构建和维护数据流管道的经验。
- 有实施数据对账、数据质量检查和验证流程的经验。
- 熟练掌握GitHub仓库管理和使用GitHub Actions进行CI/CD自动化。
- 有使用AI工具自动化工作流并提升工程效率的经验。
- 能够独立负责交付成果,并在最少监督下完成工作。
- 有主框架数据源或遗留系统集成经验者优先。
- 熟悉MongoDB或其他NoSQL数据库者优先。
COMPENSATION & BENEFITS
这是一份W2合同职位。时薪为65美元/小时。
LOCATION
该职位100%远程办公。欢迎美国任何地区的候选人申请。
查看英文原文
ABOUT THE ROLE
This is a fully hands-on individual contributor role on a data engineering team within the healthcare industry. You will own the end-to-end movement and transformation of data, from ingesting mainframe source streams into AWS through to provisioning clean, consumer-ready datasets. Your work directly enables downstream analytics and operational systems at scale.
WHAT YOU'LL DO
- Stream and process mainframe source data into AWS S3 using modern data pipeline tooling.
- Design and implement ETL/ELT workflows using AWS Glue to curate and transform data.
- Perform data reconciliation, validation, and quality checks across ingested datasets.
- Provision clean, consumer-ready datasets through AWS Aurora and RDS PostgreSQL databases.
- Manage S3 storage including data retention policies, archival strategies, and lifecycle management.
- Automate and deploy workflows using GitHub Actions and CI/CD pipelines.
- Leverage AI tools to improve engineering productivity and automate data workflows.
WHAT WE'RE LOOKING FOR
- 5 or more years of professional Data Engineering experience delivering data pipelines, ETL/ELT workflows, or data platform solutions.
- Hands-on production experience with AWS Glue for ETL/ELT design and implementation.
- Strong working knowledge of AWS S3, RDS, and Aurora in production environments.
- Demonstrated experience building and maintaining data streaming pipelines using Kafka.
- Experience implementing data reconciliation, data quality checks, and validation processes.
- Proficiency with GitHub repository management and GitHub Actions for CI/CD automation.
- Experience using AI tools to automate workflows and boost engineering productivity.
- Ability to own deliverables and drive work to completion with minimal oversight.
- Experience with mainframe data sources or legacy system integration is a plus.
- Familiarity with MongoDB or other NoSQL databases is a plus.
COMPENSATION & BENEFITS
This is a W2 contract engagement. The bill rate is $65/hour.
LOCATION
This role is 100% remote. Candidates based anywhere in the United States are welcome to apply.