远程工作雷达

数据工程师

Data Engineer

开发工程限定地区(需当地身份)
公司Pavago
薪资未公开
工作地点United States
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

职位名称:数据工程师
职位类型:全职,远程办公
工作时间:美国客户的工作时间(可根据管道监控、部署和数据刷新周期灵活调整)
职位描述
我们的客户正在寻找一名数据工程师,负责设计、构建和维护可扩展的数据基础设施以及可靠的数据显示管道,以支持整个业务的分析、报告和运营决策。
该职位需要扎实的软件工程基础,对现代数据栈有深入经验,并热衷于构建干净、可靠且高性能的数据系统。数据工程师将确保数据从源系统无缝流入数据仓库、仪表板和下游应用程序,同时保持高质量、治理和可扩展性的高标准。
理想的候选人具备分析能力,注重细节,并能够在工程、分析和业务团队之间协作,提供可信且可操作的数据。
职责
管道开发与数据集成

  • 使用 Python、SQL 或 Scala 构建、维护和优化 ETL/ELT 管道
  • • 使用 Airflow、Prefect、Dagster 或类似编排工具编排工作流
  • • 从 API、SaaS 平台、数据库、文件和流系统中提取结构化和非结构化数据
  • • 开发可扩展的连接器和自动化数据摄入流程

数据仓库与建模

  • 管理和优化云数据仓库,如 Snowflake、BigQuery 或 Redshift
  • • 使用星型和雪片型建模技术设计可扩展的模式
  • • 实现分区、聚类、索引和性能优化策略
  • • 构建干净、适用于商业智能和报告用例的数据集

数据质量、治理与可靠性

  • 实现验证检查、异常检测、日志记录和监控,以确保数据完整性
  • • 使用 dbt 或 Great Expectations 等工具实施命名规范、血缘追踪和文档标准
  • • 维护可审计的数据流程,并确保符合 GDPR、HIPAA 或行业特定要求
  • • 监控管道健康状况并主动解决故障或不一致问题

流数据与实时数据处理

  • 使用 Kafka、Kinesis、Pub/Sub 或类似平台构建和管理实时数据管道
  • • 支持低延迟数据摄入和事件驱动架构,用于时效性应用
  • • 监控流数据基础设施并优化吞吐量和可靠性
查看英文原文

Job Title: Data Engineer
Position Type: Full-Time, Remote
Working Hours: U.S. client business hours (with flexibility for pipeline monitoring, deployments, and data refresh cycles)
About the Role
Our client is seeking a Data Engineer to design, build, and maintain scalable data infrastructure and reliable data pipelines that power analytics, reporting, and operational decision-making across the business.
This role requires strong software engineering fundamentals, deep experience with modern data stacks, and a passion for building clean, reliable, and high-performance data systems. The Data Engineer will ensure data flows seamlessly from source systems into warehouses, dashboards, and downstream applications while maintaining high standards for quality, governance, and scalability.
The ideal candidate is analytical, detail-oriented, and comfortable working across engineering, analytics, and business teams to deliver trustworthy and actionable data.
Responsibilities
Pipeline Development & Data Integration

  • Build, maintain, and optimize ETL/ELT pipelines using Python, SQL, or Scala
  • • Orchestrate workflows using Airflow, Prefect, Dagster, or similar orchestration tools
  • • Ingest structured and unstructured data from APIs, SaaS platforms, databases, files, and streaming systems
  • • Develop scalable connectors and automated ingestion workflows

Data Warehousing & Modeling

  • Manage and optimize cloud data warehouses such as Snowflake, BigQuery, or Redshift
  • • Design scalable schemas using star and snowflake modeling techniques
  • • Implement partitioning, clustering, indexing, and performance optimization strategies
  • • Build clean, analytics-ready datasets for business intelligence and reporting use cases

Data Quality, Governance & Reliability

  • Implement validation checks, anomaly detection, logging, and monitoring to ensure data integrity
  • • Enforce naming conventions, lineage tracking, and documentation standards using tools such as dbt or Great Expectations
  • • Maintain audit-ready data processes and ensure compliance with GDPR, HIPAA, or industry-specific requirements
  • • Monitor pipeline health and proactively resolve failures or inconsistencies

Streaming & Real-Time Data Processing

  • Build and manage real-time data pipelines using Kafka, Kinesis, Pub/Sub, or similar platforms
  • • Support low-latency ingestion and event-driven architectures for time-sensitive applications
  • • Monitor streaming infrastructure and optimize throughput and reliability

Collaboration & Analytics Enablement

  • Partner closely with analysts, data scientists, and business stakeholders to deliver reliable datasets
  • • Support dashboard and reporting initiatives across Tableau, Looker, or Power BI
  • • Translate business requirements into scalable data solutions and models
  • • Maintain clear technical documentation for pipelines, schemas, and workflows

Infrastructure, DevOps & Automation

  • Containerize data services using Docker and manage deployments through Kubernetes when applicable
  • • Automate deployments using CI/CD pipelines such as GitHub Actions, Jenkins, or GitLab CI
  • • Manage cloud infrastructure using Terraform, CloudFormation, or similar Infrastructure-as-Code tools
  • • Continuously optimize performance, scalability, reliability, and cloud costs

What Makes You a Perfect Fit

  • Passionate about building clean, reliable, and scalable data systems
  • • Strong debugging and problem-solving mindset with high attention to detail
  • • Balance of software engineering discipline and analytical thinking
  • • Comfortable working cross-functionally with technical and non-technical stakeholders
  • • Proactive communicator who takes ownership of data quality and reliability

Required Experience & Skills

  • 3+ years of experience in Data Engineering, Back-End Engineering, or Data Infrastructure roles

• Strong proficiency in Python and SQL

  • Experience with at least one modern data warehouse (Snowflake, Redshift, BigQuery)
  • • Hands-on experience with orchestration tools such as Airflow or Prefect
  • • Strong understanding of ETL/ELT pipelines, data modeling, and data transformation workflows
  • • Familiarity with cloud platforms such as AWS, GCP, or Azure

Preferred Experience & Skills

  • Experience with dbt for data modeling and transformation management
  • • Streaming and event-driven data pipeline experience (Kafka, Kinesis, Pub/Sub)
  • • Experience with cloud-native data services such as AWS Glue, GCP Dataflow, or Azure Data Factory
  • • Familiarity with Docker, Kubernetes, Terraform, or CI/CD workflows
  • • Background in regulated industries such as healthcare, fintech, or enterprise SaaS
  • • Experience optimizing warehouse costs and query performance at scale

What Does a Typical Day Look Like?
A Data Engineer’s day revolves around maintaining reliable pipelines, improving data quality, and enabling teams with scalable access to trustworthy data. You will:

  • Monitor pipeline health and troubleshoot failed jobs in Airflow or related orchestration systems
  • • Build and maintain ingestion pipelines for APIs, SaaS platforms, and operational databases
  • • Optimize SQL queries and warehouse performance to improve efficiency and reduce cloud costs
  • • Collaborate with analysts and data scientists to provide curated datasets for reporting and modeling
  • • Implement validation checks and monitoring to prevent downstream data quality issues
  • • Document data models, transformations, and workflows to ensure scalability and maintainability

In essence: you ensure the organization has accurate, timely, and reliable data powering operational, analytical, and strategic decisions.
Key Metrics for Success (KPIs)
• Pipeline uptime ≥ 99%

  • Data freshness maintained within agreed SLAs
  • • Zero critical data quality issues reaching downstream reporting systems
  • • Improved warehouse query performance and cost optimization
  • • Timely delivery of scalable and reliable datasets
  • • Positive feedback from analysts, data scientists, and business stakeholders

Interview Process

  • Initial Phone Screen

• Video Interview with Pavago Recruiter

  • Technical Assessment (e.g., build a small ETL pipeline or optimize a SQL query)
  • • Client Interview with Engineering/Data Team

• Offer & Background Verification
#DataEngineer #ETL #DataPipelines #BigQuery #Snowflake #Redshift #Airflow #Python #SQL #CloudData #AnalyticsEngineering #DataInfrastructure #RemoteWork #DataEngineeringJobs
Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

SEO专员

PavagoUnited StatesFull Time今天
市场运营限定地区(需当地身份)

收款专员

PavagoUnited StatesFull Time今天
其他限定地区(需当地身份)

财务建模专员

PavagoUnited StatesFull Time今天
职能支持限定地区(需当地身份)

← 返回全部职位