数据与机器学习实习生
Data and Machine Learning Intern
**Duration:** 六个月
**Format:** 全职(每周40小时),有薪
在过去一年中,Loka的团队为各种公司推出了近200个GenAI项目,包括全球排名第一的GenAI阅读导师、一个将房屋转变为电池的初创公司以及领先的抗癌实验室。而且我们还享受每个星期五放假 😎
作为数据-机器学习实习生,你将获得支持Loka认证专家、技术专家和博士的专业经验,同时提升你的技能,建立作品集并启动让你自豪的项目。
**职位职责**
- 协助设计、开发和维护数据管道,确保数据干净、可靠且及时。
- 与团队合作,实现和优化ETL流程。
- 将数据从不同来源整合到数据仓库、数据湖和湖仓中。
- 支持数据管理任务,包括数据清洗、验证和转换。
- 理解业务目标,并开发有助于实现这些目标的模型,以及跟踪进度的指标。
- 使用经典机器学习、深度学习和基础模型实现ML系统,遵循最佳实践。
- 参与客户沟通,帮助收集需求并传达交付成果。
- 仔细观察数据问题,进行数据可视化和探索,以及分析可能影响部署后性能的数据分布差异。
- 识别和分析模型错误。
**所需硬技能**
- 计算机科学或相关专业的本科最后一年学生
- 英语流利
- 熟悉Python、机器学习和数据相关库
- 熟悉数据库
- 理解统计学、机器学习和深度学习算法
- 有可视化和操作大数据集的经验
- 问题解决能力
- 附加:熟悉AWS、(Py)Spark、Airflow、数据湖和数据仓库
**所需软技能**
- 好奇心:你渴望在不同行业中利用现代技术栈学习和成长。
- 自主性和积极态度:我们是一个完全远程、全球分布的团队。
- 团队合作:喜欢协作的方式。
- 适应性:以创业心态运作,保持创业节奏。
- 可靠性:你可以被信任完成高质量的工作。
**福利**
- 每隔一周的星期五休息
- 健康补贴
- 远程和灵活工作
- 有薪病假和当地节假日
**请提交英文简历。**
查看英文原文
**Duration:** Six months
**Format:** Full time (40 hrs/week), paid
In the last year at Loka, our teams launched almost 200 GenAI projects for companies of all kinds, including the world’s Number 1 GenAI reading tutor, a startup that transforms homes into batteries and a leading cancer-fighting laboratory. And we did it all while enjoying every other Friday off 😎
As a Data - Machine Learning Intern, you'll gain professional experience supporting Loka’s certified specialists, technical experts and PhDs while elevating your skillset, building a portfolio and launching projects you’re proud of.
**The Role**
- Assist in designing, developing and maintaining data pipelines to ensure clean, reliable and timely data.
- Collaborate with the team to implement and optimize ETL processes.
- Integrate data from various sources into warehouses, data lakes and lakehouses.
- Support data management tasks, including data cleaning, validation and transformation.
- Understand business objectives and develop models that help achieve them, plus metrics to track their progress.
- Implement ML systems using classical ML, DL and Foundation Models following best practices.
- Participate in client communications by helping gather requirements and communicate deliverables.
- Explore and visualize data with a careful eye for issues that require data cleaning as well as differences in data distribution that may affect performance after deployment.
- Identify and analyze model errors.
**Required Hard Skills**
- Last year of a bachelor’s degree in Computer Science or related
- Proficient in English
- Basic knowledge of Python, ML, and Data libraries
- Basic knowledge of Databases
- Understanding of statistical, ML and deep learning algorithms
- Experience visualizing and manipulating big datasets
- Problem solving
- Bonus: AWS knowledge, (Py)Spark, Airflow, Data Lakes and Data Warehouses
**Required Soft Skills**
- Curiosity: You’re ambitious to learn and grow in different industries utilizing a modern tech stack.
- Autonomy and positivity: We’re a fully remote, globally distributed team.
- Teamwork: Enjoy a collaborative approach.
- Adaptability: Operate with a startup mindset and move at a startup pace.
- Dependable: You can be trusted to deliver high-quality work.
**Benefits**
- Every other Friday off
- Health Bonus
- Remote and flexible
- Paid sick days and local holidays
**Please submit your CV in English.**