高级数据工程师
Senior Data Engineer
- 编写并验证针对大规模生产数据集的诊断SQL查询
- 构建并维护将竞价、成交和展示日志导入BigQuery的摄入管道
- 协调独立设计的数据集中的字段,并维护版本化的字段映射
- 开发时点正确的特征表和聚合管道
- 设计并维护带有延迟标签处理的转化和标注管道
- 负责数据服务的写入路径、模式契约、发布流程和新鲜度SLO
- 构建实验基础设施,包括流量分割和报告管道
- 执行大规模历史回填,并在映射变更后安全地重新处理数据
- 实现数据隔离和安全聚合控制,以保护广告客户数据
- 开发自动化的数据质量验证框架
- 与客户工程师紧密合作,并准备操作文档
- 参与架构讨论并推动平台可扩展性改进
- 5年以上数据工程经验
- 至少2年使用生产环境机器学习或大规模分析管道的经验
- 精通SQL,包括窗口函数和增量处理模式
- 强大的Python技能,用于生产级管道开发
- 具有Spark或PySpark的实际经验
- 具有使用Airflow、Cloud Composer、Dagster或其他类似工具设计ETL/ELT管道的经验
- 具有在云数据仓库(尤其是BigQuery)上大规模工作的经验
- 对数据建模和时点正确性有深刻理解
- 具有处理大规模事件驱动或点击流数据集的经验
- 具有支持业务关键生产管道的经验
- 英语水平为中高级或以上
额外加分项
- 具有GCP服务(包括Dataflow、Pub/Sub、GCS和Beam)的经验
- 具有构建流式或近实时摄入系统的经验
- 了解特征存储、训练/推理偏差和标签泄露预防
- 具有广告技术或基于拍卖的环境经验
- 具有处理机器学习系统中延迟或不完整标签的经验
- 具有dbt或其他类似转换框架的经验
- 具有将解决方案部署到客户自有基础设施的经验
- 了解与GDPR/CCPA相关的隐私工程实践
- 具有实验基础设施和统计验证管道的经验
- 具有在混合云/本地Linux环境中工作的经验
查看英文原文
- Write and defend diagnostic SQL queries against large-scale production datasets
- Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
- Harmonize fields across independently designed datasets and maintain versioned field mappings
- Develop point-in-time-correct feature tables and aggregation pipelines
- Design and maintain conversion and labeling pipelines with delayed label handling
- Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
- Build experimentation infrastructure including traffic splitting and reporting pipelines
- Perform large-scale historical backfills and safe reprocessing after mapping changes
- Implement data isolation and safe-aggregation controls for advertiser data protection
- Develop automated data quality validation frameworks
- Collaborate closely with Customer engineers and prepare operational documentation
- Contribute to architecture discussions and platform scalability improvements
- 5+ years of experience in Data Engineering
- At least 2 years of experience working with production ML or large-scale analytics pipelines
- Expert-level SQL skills including window functions and incremental processing patterns
- Strong Python skills for production-grade pipeline development
- Hands-on experience with Spark or PySpark
- Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
- Experience working with cloud data warehouses at scale, preferably BigQuery
- Strong understanding of data modeling and point-in-time correctness
- Experience working with event-driven or clickstream datasets at very large scale
- Experience supporting business-critical production pipelines
- Upper-Intermediate English level or higher
WILL BE A PLUS
- Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
- Experience building streaming or near-real-time ingestion systems
- Understanding of feature stores, train/serve skew, and label leakage prevention
- Experience in AdTech or auction-based environments
- Experience handling delayed or incomplete labels in ML systems
- Experience with dbt or similar transformation frameworks
- Experience delivering solutions into Customer-owned infrastructure
- Knowledge of GDPR/CCPA-related privacy engineering practices
- Experience with experimentation infrastructure and statistical validation pipelines
- Experience working in hybrid cloud/on-prem Linux environments
- Terraform and Kubernetes experience
- Experience optimizing warehouse cost and performance
PERSONAL PROFILE
- Strong analytical and problem-solving skills
- Ownership-oriented mindset
- Ability to work independently in a client-facing environment
- Strong communication and documentation skills
- Comfortable working in a fast-paced engineering environment
- Collaborative and proactive attitude
Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. We are looking for a Senior Data Engineer who enjoys solving complex distributed data challenges and building production-grade ML-oriented data systems.
You will become part of a dedicated Sigma Software team developing predictive modeling and optimization capabilities for a live advertising ecosystem. The role combines large-scale event processing, streaming and batch pipelines, experimentation infrastructure, and high-throughput data engineering in a cloud-native environment.
We as a company offer the opportunity to work on impactful global products, collaborate with experienced engineers, and contribute to architecture decisions while growing your expertise in large-scale distributed systems and modern data platforms.
CUSTOMER
Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a large-scale ad exchange handling hundreds of millions of auction requests per day and is actively investing in predictive decisioning technologies to optimize advertising outcomes in real time.
PROJECT
The project focuses on building a predictive modeling and optimization platform on top of a live ad exchange environment. The platform performs real-time supply scoring and filtering, contextual performance estimation, look-alike audience generation, and multi-objective optimization under business constraints.
The solution processes massive-scale event and auction datasets and includes feature engineering pipelines, streaming and batch ingestion, experimentation infrastructure, point-in-time-correct training data generation, and ML-oriented data services with strict operational reliability and compliance requirements.
Originally posted on Himalayas