数据工程师
Data Engineer
职位概述
我们正在寻找一位经验丰富的数据工程师,在Google云平台(GCP)上设计、构建和维护可扩展的元数据驱动的数据摄入和转换框架。该职位将负责开发配置驱动的数据摄入流水线,接入多个源系统,并实现原始层、青铜层、银层和金层数据。数据工程师将使用BigQuery、Cloud Composer/Apache Airflow、Python、Dataform、Datastream、Pub/Sub、Dataflow、Cloud Run和Cloud Storage等技术,提供可靠且可重用的数据解决方案。理想的候选人应具备元数据驱动的流水线开发、CDC、高级SQL、Python、数据质量、对账、自动化测试和CI/CD方面的丰富经验。工程师将与数据治理和其他平台相关方紧密合作,确保数据流水线具备可扩展性、受控性、可审计性、可测试性和生产就绪性。
职责
- 使用可重用的流水线模板设计和构建元数据驱动和配置驱动的数据摄入框架。
- 设计、开发和维护涵盖源系统、源对象、加载配置、执行参数、运行日志、审计信息和重新处理的元数据配置模型。
- 构建动态的Cloud Composer / Apache Airflow DAG,根据元数据配置生成和执行数据摄入工作流。
- 开发可重用的数据摄入模式和流水线生成器,而不是为每个源表单独构建流水线。
- 接入多个源系统,包括关系型数据库、SAP、REST API、文件、流式数据源和其他企业平台。
- 按照定义的湖仓架构和工程标准,实现原始层和青铜层的数据摄入。
- 实现源到目标的对账,以确保摄入数据的完整性和准确性。
- 实现强大的操作控制,包括错误处理、重试、隔离处理、警报、监控和审计日志。
- 开发按批次重放和重新处理功能,以支持失败的加载、历史重新加载和受控的数据恢复。
- 支持批量、增量、流式和变更数据捕获(CDC)的数据摄入模式。
- 实现青铜层处理,包括数据类型、清洗、模式验证和模式强制。
- 使用适当的源键、主键或业务键实现数据去重。
查看英文原文
Job Summary
We are looking for an experienced Data Engineer to design, build, and maintain a scalable, metadata-driven data ingestion and transformation framework on Google Cloud Platform (GCP).
The role will be responsible for developing configuration-driven ingestion pipelines, onboarding multiple source systems, and implementing Raw, Bronze, Silver, and Gold data layers. The Data Engineer will use technologies including BigQuery, Cloud Composer/Apache Airflow, Python, Dataform, Datastream, Pub/Sub, Dataflow, Cloud Run, and Cloud Storage to deliver reliable and reusable data solutions.
The ideal candidate will have strong experience in metadata-driven pipeline development, CDC, advanced SQL, Python, data quality, reconciliation, automated testing, and CI/CD. The engineer will work closely with Data Governance and other platform stakeholders to ensure that data pipelines are scalable, governed, auditable, testable, and production-ready.
Responsibilities
- Design and build a metadata-driven and configuration-driven ingestion framework using reusable pipeline templates.
- Design, develop, and maintain the metadata configuration model covering source systems, source objects, load configurations, execution parameters, run logs, audit information, and reprocessing.
- Build dynamic Cloud Composer / Apache Airflow DAGs that generate and execute ingestion workflows based on metadata configurations.
- Develop reusable ingestion patterns and pipeline generators instead of building separate pipelines for individual source tables.
- Onboard multiple source systems, including relational databases, SAP, REST APIs, files, streaming sources, and other enterprise platforms.
- Implement data ingestion into Raw and Bronze layers following defined lakehouse architecture and engineering standards.
- Implement source-to-target reconciliation to ensure completeness and accuracy of ingested data.
- Implement robust operational controls, including error handling, retries, quarantine processing, alerting, monitoring, and audit logging.
- Develop replay-by-batch and reprocessing capabilities to support failed loads, historical reloads, and controlled data recovery.
- Support batch, incremental, streaming, and Change Data Capture (CDC) ingestion patterns.
- Implement Bronze-layer processing, including data typing, cleansing, schema validation, and schema enforcement.
- Implement data de-duplication using appropriate source keys, primary keys, or business keys.
- Implement soft-delete representation and appropriate handling of deleted source records.
- Implement CDC change-history materialization and maintain historical changes where required.
- Develop appropriate BigQuery partitioning, clustering, MERGE, and incremental processing strategies.
- Perform data compaction and other performance optimization activities where required.
- Design and develop Silver and Gold transformation models using Dataform.
- Develop reusable transformation components, macros, dependencies, and incremental transformation strategies.
- Implement assertions and automated tests for transformation models and ensure no untested transformation reaches production.
- Implement in-pipeline data quality checks, validation rules, and automated promotion gates.
- Work closely with the Data Governance Consultant to incorporate data quality, governance, lineage, audit, and control requirements into the engineering framework.
- Prevent data that fails critical quality requirements from being promoted to downstream layers.
- Implement geospatial data ingestion, including GeoJSON processing and conversion to BigQuery GEOGRAPHY.
- Ensure appropriate preservation and handling of Spatial Reference System (SRS) information for geospatial datasets.
- Implement ingestion and management of semi-structured and unstructured data using Google Cloud Storage and BigQuery object tables.
- Follow Git-based development practices, including branching, pull/merge requests, peer reviews, and code reviews.
- Develop automated tests as a standard part of pipeline and transformation development.
- Integrate data pipelines and Dataform transformations with CI/CD processes.
- Document each source-system onboarding, including configuration, mappings, dependencies, reconciliation, operational procedures, and troubleshooting guidance.
- Develop and maintain operational runbooks and onboarding documentation that enable ESNAD teams to independently onboard additional data sources.
Must-Have Skills
- Minimum 5+ years of professional Data Engineering experience.
- Minimum 2+ years of hands-on experience building data pipelines on Google Cloud Platform (GCP).
- Strong hands-on experience with Google Cloud Platform data engineering services.
- Expert-level SQL skills, including complex analytical SQL development.
- Strong hands-on experience with Google BigQuery, including data modeling and large-volume data processing.
- Strong Python programming skills for data engineering, automation, API integration, validation, and framework development.
- Strong hands-on experience with Cloud Composer / Apache Airflow.
- Proven experience developing dynamic Airflow DAGs.
- Experience developing metadata/configuration-driven DAG generation.
- Experience with Airflow sensors, custom operators, task dependencies, scheduling, retries, and failure management.
- Proven experience designing and implementing metadata-driven or configuration-driven ingestion frameworks.
- Experience developing pipeline generators and reusable ingestion templates, rather than only building individual pipelines for individual tables.
- Experience designing metadata configurations for source systems, source objects, load parameters, runtime configurations, audit logs, and reprocessing.
- Hands-on experience with Google Cloud Datastream or an equivalent log-based CDC technology.
- Strong understanding of CDC patterns, including inserts, updates, deletes, soft deletes, incremental loads, and historical change management.
- Experience implementing CDC recovery, reconciliation, replay, and reprocessing patterns.
- Hands-on experience with Dataform or dbt.
- Strong understanding of transformation model structures, macros, reusable logic, assertions, automated tests, dependencies, and incremental strategies.
- Experience developing Raw, Bronze, Silver, and Gold data layers or equivalent medallion/lakehouse architecture.
- Experience implementing source-to-target data reconciliation and validation.
- Experience implementing exception handling, quarantine processes, retry mechanisms, operational monitoring, and alerting.
- Working knowledge of Google Cloud Pub/Sub and event-driven or streaming ingestion patterns.
- Working knowledge of Dataflow / Apache Beam for data processing.
- Working knowledge of Cloud Run and/or Cloud Data Fusion.
- Strong understanding of relational database extraction and database-engine-specific CDC constraints.
- Working knowledge of SAP data extraction patterns, including ODP, SLT, or certified SAP connectors.
- Experience with REST API ingestion, including authentication, pagination, error handling, and incremental extraction.
- Experience with file-based data ingestion and file manifest validation.
- Experience implementing automated data quality checks, validation rules, and pipeline quality gates.
- Strong knowledge of BigQuery partitioning, clustering, MERGE operations, de-duplication, schema enforcement, and incremental processing.
- Strong understanding of data pipeline monitoring, auditability, traceability, and operational support.
- Strong experience with Git-based development and version control.
- Experience working with peer reviews, code reviews, and controlled development workflows.
- Experience writing unit, integration, data quality, or pipeline tests as a standard engineering practice.
- Experience implementing or working with CI/CD pipelines for data engineering workloads.
- Ability to produce high-quality technical documentation, source onboarding documentation, and operational runbooks.
Good-to-Have Skills
- Google Cloud Professional Data Engineer certification.
- Google Cloud Associate Cloud Engineer certification.
- dbt Analytics Engineering Certification.
- Advanced hands-on experience with Google Cloud Datastream and enterprise-scale CDC implementations.
- Advanced experience with Dataflow / Apache Beam for batch and streaming workloads.
- Advanced experience with Pub/Sub and event-driven data architectures.
- Advanced BigQuery performance optimization and cost-optimization experience.
- Experience designing and implementing enterprise-scale lakehouse or medallion architectures.
- Experience designing reusable enterprise data ingestion frameworks and platform accelerators.
- Advanced hands-on experience with Dataform, including reusable components, assertions, incremental models, and CI/CD.
- Experience designing enterprise-level automated data quality frameworks.
- Experience with data governance, metadata management, data lineage, data cataloging, auditability, and traceability.
- Strong hands-on experience with SAP ODP, SAP SLT, or certified SAP extraction connectors.
- Experience implementing SAP CDC and incremental extraction patterns.
- Experience processing GeoJSON and BigQuery GEOGRAPHY data.
- Knowledge of geospatial data engineering and Spatial Reference Systems (SRS).
- Experience handling semi-structured and unstructured data at enterprise scale.
- Experience with Google Cloud Storage and BigQuery object tables.
- Experience implementing CI/CD pipelines for Dataform, dbt, Airflow, and other data engineering workloads.
- Experience with automated deployments, environment promotion, testing gates, and release-management practices.
- Knowledge of cloud security, IAM, service accounts, secrets management, and secure data pipeline design.
- Experience implementing monitoring, observability, logging, alerting, and operational dashboards for enterprise data pipelines.
- Experience working in large enterprise data transformation or cloud migration programs.
- Experience working closely with Data Governance, Data Architecture, Security, DevOps, and Business teams.
- Experience mentoring junior and mid-level Data Engineers and establishing engineering standards.
- For a Senior Data Engineer / Technical Lead, 8+ years of overall Data Engineering experience is preferred, along with demonstrated technical leadership and architecture ownership.
Originally posted on Himalayas