数据工程师
Data Engineer
职位概述
我们正在寻找一位经验丰富且多才多艺的数据工程师加入我们充满活力、快速发展的团队。如果你对数据充满热情,擅长解决复杂问题,并能直接与企业级客户合作,将业务需求转化为可扩展的技术解决方案,那么这个职位可能是你的理想选择。
ShyftLabs 是一家成立于 2020 年初的数据产品公司,主要为财富 500 强企业提供服务。我们通过创新创造价值,为多个行业的企业提供数字化解决方案,助力企业加速增长。
除了扎实的技术能力,我们还希望你具备较强的企业意识和与客户及利益相关者沟通的能力。理想的候选人应能够与企业级客户协作,将复杂的概念转化为业务成果,并确保工程执行与战略目标保持一致。
岗位职责
- 使用云服务如 GCP Dataflow、Cloud Functions、Pub/Sub 和 Cloud Composer 设计、构建和维护可扩展且可靠的批处理和实时 ETL/ELT 数据流水线。
- 架构并实现能够处理高容量数据摄入和处理的稳健数据基础设施。
- 开发和管理我们的中心数据仓库 Google BigQuery。
- 设计和实现优化性能、可扩展性和长期可维护性的数据模型、模式和表结构。
- 编写干净、高效且易于维护的 SQL 和 Python 代码,将原始数据转换为经过整理、可供分析的数据集。
- 构建可靠的转换工作流,支持分析、报告和数据科学项目。
- 监控、排查和优化数据基础设施,以确保高性能、可靠性和成本效率。
- 实施 BigQuery 最佳实践,包括分区、聚类、查询优化和物化视图。
- 构建和维护作为商业智能和报告“唯一真实来源”的定制数据模型。
- 确保数据经过优化,便于 BI 工具如 Looker 和其他分析平台使用。
- 实现自动化数据质量检查、验证规则和监控框架,以确保数据流水线和仓库系统的完整性和可靠性。
- 建立数据治理、可观测性和数据血缘追踪的流程。
- 与跨职能团队紧密合作,确保数据架构符合业务需求和技术标准。
查看英文原文
Position Overview
We are looking for an experienced and versatile Data Engineer to join our dynamic and fast-growing team. If you are passionate about data, solving complex problems, and working directly with enterprise stakeholders to translate business needs into scalable technical solutions, this role could be the perfect fit.
ShyftLabs is a growing data product company that was founded in early 2020 and works primarily with Fortune 500 companies. We deliver digital solutions built to help accelerate the growth of businesses across various industries by focusing on creating value through innovation.
In addition to strong technical expertise, we are seeking someone with strong business awareness and the ability to lead client and stakeholder communication. The ideal candidate will be comfortable collaborating with enterprise-level clients, translating complex technical concepts into business outcomes, and ensuring alignment between engineering execution and strategic objectives.
Why You’ll Love Working at ShyftLabs
At ShyftLabs, your work matters. We’re a growing data product company making a big impact with Fortune 500 clients and as we scale, you’ll have the chance to shape solutions, influence strategy, and grow your career alongside us.
Here’s what you can expect when you join our team:
-Work Arrangement: This role is currently fully remote, providing flexibility to work from home. As the team and organization continue to grow, there may be an opportunity for the role to transition into a hybrid work model in the future, with occasional in-office collaboration.
-Comprehensive Benefits: We cover 100% of health, dental, and vision insurance premiums for you and your dependents which means no out-of-pocket costs. Eligibility starts from day one itself.
-Growth & Learning: Access extensive learning and development resources to keep leveling up your skills.
Inclusion at ShyftLabs
We’re building something big, and we want you on the journey with us. If you’re ready to use data and innovation to make an impact, apply today and let’s grow together.
ShyftLabs is an equal-opportunity employer committed to creating a safe, diverse, and inclusive environment. We encourage applicants of all backgrounds including ethnicity, religion, disability status, gender identity, sexual orientation, family status, age, and nationality to apply. If you require accommodation during the interview process, let us know and we’ll be happy to support you.
Job Responsibilities
- Design, build, and maintain scalable and reliable batch and real-time ETL/ELT data pipelines using cloud services such as GCP Dataflow, Cloud Functions, Pub/Sub, and Cloud Composer.
- Architect and implement robust data infrastructure capable of handling high-volume data ingestion and processing.
- Develop and manage our central data warehouse in Google BigQuery.
- Design and implement data models, schemas, and table structures optimized for performance, scalability, and long-term maintainability.
- Write clean, efficient, and maintainable SQL and Python code to transform raw data into curated, analysis-ready datasets.
- Build reliable transformation workflows that support analytics, reporting, and data science initiatives.
- Monitor, troubleshoot, and optimize data infrastructure to ensure high performance, reliability, and cost efficiency.
- Implement BigQuery best practices, including partitioning, clustering, query optimization, and materialized views.
- Build and maintain curated data models that serve as the “source of truth” for business intelligence and reporting.
- Ensure data is optimized and readily accessible for BI tools such as Looker and other analytics platforms.
- Implement automated data quality checks, validation rules, and monitoring frameworks to ensure the integrity and reliability of data pipelines and warehouse systems.
- Establish processes for data governance, observability, and lineage tracking.
- Work closely with software engineers, data analysts, and data scientists to understand their data requirements and provide the necessary infrastructure and data products.
- Lead and support client and stakeholder communication, working with enterprise clients to translate business needs into scalable data solutions.
- Partner with product teams and leadership to ensure that technical data solutions align with business strategy and client expectations.
- Take ownership of data platforms and architecture decisions, helping shape the future direction of our analytics and data infrastructure.
- Identify opportunities to improve data reliability, automate workflows, and generate new insights through data.
- Contribute to a collaborative, high-performing engineering culture with strong communication and teamwork.
Basic Qualifications
- 5+ years of hands-on experience in data engineering, data integration, or data platform development.
- Degree in Computer Science, Engineering, Mathematics, or related STEM discipline.
- Strong programming and query skills in SQL and Python.
- Experience working with distributed version control systems such as Git in an Agile/Scrum environment.
- Experience designing and orchestrating ETL pipelines, particularly with Databricks.
- Experience working within cloud environments (GCP, AWS, or Azure).
- Experience with database systems such as MongoDB and Elasticsearch.
- Strong understanding of data warehousing and dimensional modeling methodologies.
- Hands-on experience with Airflow and Hadoop.
- Experience using Docker for containerized workflows and reproducible environments.
- Ability to identify opportunities to improve data quality, reliability, and automation.
- Strong business awareness and communication skills, with the ability to collaborate with both technical teams and business stakeholders.
- Experience within the retail industry is a plus.
Preferred Qualifications
- Master’s degree in Computer Science, Engineering, or related discipline.
- Experience working with enterprise-scale data platforms and Fortune 500 clients.
- Familiarity with Druid and its Python API, including Kafka integrations.
- Strong experience using Apache Spark for large-scale data processing.
- Experience designing real-time streaming data architectures.
- Experience working with AI-driven platforms, data infrastructure supporting AI/ML systems, or agentic AI workflows