数据工程师
Data Engineer
这是一个远程职位。
我们正在寻找一名具有大约5年经验的数据工程师,负责设计、构建和维护可扩展的数据平台和管道。理想的候选人应具备现代数据工程实践、分布式数据处理技术和开放数据湖架构的深厚专业知识。该职位需要与跨职能团队紧密合作,确保可靠、高效且高质量的数据解决方案,以支持分析和业务需求。主要职责包括:
· 使用银牌架构原则构建可靠的数据显示管道。
- 管理和优化开放存储格式,包括Parquet、Iceberg和Delta Lake。
- 使用Databricks和Apache Spark设计、开发和维护分布式数据处理工作负载。
- 执行PostgreSQL和SQLite环境中的数据库操作、优化和维护。
- 确保 across 数据平台的数据质量、可靠性、可扩展性和性能。
- 与数据使用者和利益相关者合作,支持报告、分析和运营数据需求。
- 监控、排查并改进数据基础设施和管道性能。
- 遵循数据工程最佳实践、编码标准和文档流程。
要求
- 计算机科学、软件工程或相关领域的学士学位。
- 大约5年数据工程或相关领域的工作经验。
- 精通SQL,特别是PostgreSQL。
- 熟练使用Python和DuckDB。
- 具有数据湖和开源查询与存储层的经验。
- 有使用Parquet、Iceberg和Delta Lake的实际经验。
- 有使用分布式计算框架的经验,特别是Databricks和Apache Spark。
- 对数据建模、ETL/ELT流程和数据管道开发有深刻理解。
- 了解数据库管理、性能调优和优化技术。
- 优秀的解决问题和分析能力。
- 能够独立工作,并在协作团队环境中有效工作。
- 优秀的英语沟通能力。
此职位对埃及申请人是远程办公,对黎巴嫩申请人是混合办公。
最初发布于Himalayas
查看英文原文
This is a remote position.
We are seeking a Data Engineer with approximately 5 years of experience to design, build, and maintain scalable data platforms and pipelines. The ideal candidate will have strong expertise in modern data engineering practices, distributed data processing technologies, and open data lake architectures. This role involves working closely with cross-functional teams to ensure reliable, efficient, and high-quality data solutions that support analytics and business needs. The Key Responsibilities are:
· Build resilient data pipelines using Medallion architecture principles.
- Manage and optimize open storage formats including Parquet, Iceberg, and Delta Lake.
- Design, develop, and maintain distributed data processing workloads using Databricks and Apache Spark.
- Perform database operations, optimization, and maintenance for PostgreSQL and SQLite environments.
- Ensure data quality, reliability, scalability, and performance across data platforms.
- Collaborate with data consumers and stakeholders to support reporting, analytics, and operational data requirements.
- Monitor, troubleshoot, and improve data infrastructure and pipeline performance.
- Follow data engineering best practices, coding standards, and documentation processes.
Requirements
Requirements
· Bachelor’s degree in Computer Science, Software Engineering, or a related field.
- Approximately 5 years of experience in Data Engineering or a related role.
- Advanced SQL skills with strong expertise in PostgreSQL.
- Strong proficiency in Python and DuckDB.
- Experience with data lakes and open-source query and storage layers.
- Hands-on experience working with Parquet, Iceberg, and Delta Lake.
- Experience with distributed computing frameworks, specifically Databricks and Apache Spark.
- Strong understanding of data modeling, ETL/ELT processes, and data pipeline development.
- Knowledge of database administration, performance tuning, and optimization techniques.
- Excellent problem-solving and analytical skills.
- Ability to work effectively both independently and within a collaborative team environment.
- Excellent English communication skills.
This opportunity is remote for Egyptian applicants and Hybrid for Lebanese applicants.
Originally posted on Himalayas