数据工程实习生
Data Engineering Intern
**时长:** 三个月
**形式:** 全职(每周40小时),有薪
在过去一年中,Loka的团队为各种公司推出了近200个GenAI项目,包括全球排名第一的GenAI阅读导师、一个将房屋转变为电池的初创公司以及一家领先的抗癌实验室。2024年底,Loka被AWS评为年度创新合作伙伴,从15万个合作伙伴中脱颖而出。而我们还享受每星期五的假期 😎。
作为数据工程实习生,你将获得实际的专业经验,支持Loka的认证专家、技术专家和博士,同时提升你的技能,建立作品集并启动让你自豪的项目。
##### 你将负责
- 协助设计、开发和维护数据管道,确保数据干净、可靠且及时。
- 与团队合作实施和优化ETL管道。
- 将数据从各种来源整合到数据仓库、数据湖和湖仓中。
- 支持数据管理任务,包括数据清洗、验证和转换。
- 理解业务目标,并帮助开发支持这些目标的数据模型,以及跟踪进展的指标。
- 参与客户沟通,协助收集需求并传达交付成果。
- 以细致的眼光探索和可视化数据,识别需要清洗的问题,以及可能影响部署后性能的数据分布差异。
- 识别数据质量问题,并寻找改进代码库的机会。
##### 你将带来的
**经验与技术技能**
- 精通英语。
- 基础的Python和数据库知识。
- 基础的SQL和数据库引擎知识。
- 有可视化和操作数据集的经验。
- 强大的问题解决能力。
- 附加优势:熟悉AWS,(Py)Spark,Airflow,数据湖和数据仓库,git。
##### 额外要求
- 英语优秀,我们所有会议、客户电话和商务沟通都使用英语。
- 提交的简历需为英文版。
##### 个性特征
- **好奇:** 你渴望通过现代技术栈在不同行业中学习和成长。
- **自主积极:** 你在完全远程、全球分布的团队中表现出色。
- **团队合作:** 你喜欢协作的方式。
- **适应力强:** 你具备初创公司的思维和节奏。
- **可靠:** 你可以被信任完成高质量的工作。
查看英文原文
**Duration:** Three months
**Format:** Full time (40 hrs/week), paid
In the last year at Loka, our teams launched almost 200 GenAI projects for companies of all kinds, including the world's number 1 GenAI reading tutor, a startup that transforms homes into batteries and a leading cancer fighting laboratory. To cap it off, at the end of 2024 Loka was recognized by AWS as Innovation Partner of the Year, outshining 150,000 partners for the title. And we did it all while enjoying every other Friday off 😎.
As a Data Engineering Intern, you'll gain hands on professional experience supporting Loka's certified specialists, technical experts and PhDs, all while elevating your skillset, building a portfolio and launching projects you're proud of.
##### What You'll Do
- Assist in designing, developing and maintaining data pipelines to ensure clean, reliable and timely data.
- Collaborate with the team to implement and optimize ETL pipelines.
- Integrate data from various sources into warehouses, data lakes and lakehouses.
- Support data management tasks, including data cleaning, validation and transformation.
- Understand business objectives and help develop data models that support them, along with metrics to track progress.
- Participate in client communications by helping gather requirements and communicate deliverables.
- Explore and visualize data with a careful eye for issues that require cleaning, as well as differences in data distribution that may affect performance after deployment.
- Identify data quality issues and opportunities to improve the codebase.
##### What You'll Bring
**Experience & Technical Skills**
- Proficient in English.
- Basic knowledge of Python and data libraries.
- Basic knowledge of SQL and database engines.
- Experience visualizing and manipulating datasets.
- Strong problem solving skills.
- Bonus: AWS knowledge, (Py)Spark, Airflow, data lakes and data warehouses, git.
##### Additional Requirements
- Excellent English, we work entirely in English for meetings, client calls and business communications.
- CV submitted in English.
##### Personality Profile
- **Curious:** You're ambitious to learn and grow in different industries using a modern tech stack.
- **Autonomous and positive:** You excel in a fully remote, globally distributed team.
- **Team player:** You enjoy a collaborative approach.
- **Adaptable:** You operate with a startup mindset and move at a startup pace.
- **Dependable:** You can be trusted to deliver high quality work.
##### Benefits
- Every other Friday off
- Remote and flexible
- Paid sick days and local holidays
- Fitness subscription
**_Your achievements matter to us! Ensure your CV and LinkedIn profile are up to date and accurately reflect your experience._**