机器学习基础设施工程师
ML Infra Engineer
关于 Humble Robotics
在 Humble Robotics 工作意味着参与几十年来地面运输领域最大的变革。我们正在打造一款自动驾驶、零排放的运输车,通过突破性的基于视觉的 AI 技术大幅降低货运成本,专为当今全球物流网络设计。
我们是一支快速前进、紧密协作的团队,成员来自 AV 行业资深人士和富有创新思维的人。我们认为文化无法被设计出来——但当它自然形成时,这将是一次终生难忘的冒险。
进步从未如此贴近现实。
职位概述
我们正在寻找一名 ML 基础设施工程师,帮助设计、构建和扩展实现我们雄心勃勃愿景所需的基础系统。你将参与支持 ML 训练循环每个阶段的工具和基础设施,并在塑造我们工作的技术与组织决策中发挥重要作用。从车辆计算到数据收集,再到数据集整理,以及大规模模型训练和部署,帮助我们构建可靠、高效且安全的基础设施,让 Humble Robotics 的每个团队都能依赖它。这里很有趣。我们在做酷的事情。
理想的候选人是一位善于从基础原理思考问题的人,能够胜任广泛领域的通用角色。参与每一层技术栈,帮助让软件迭代循环尽可能快速和高效。我们是一个小团队,你的意见、经验和知识将在塑造我们构建、运营并依赖的每一个系统中发挥关键作用。
主要职责
- 参与数据收集基础设施的开发,确保传感器数据能可靠高效地从我们的车辆传输到 ML 平台
- 开发批处理计算管道,用于对原始数据进行目录管理、探索和整理,生成高质量的训练集
- 设计并扩展在 GPU 集群上的分布式 ML 训练
- 对整个流程的性能、可观测性、效率和安全性负责
- 与 ML 团队合作,理解他们的工作流程,并将其转化为能加速他们工作的可靠基础设施
最低要求
- 在云基础设施上构建和运维高可用性 Web 服务的经验
- 使用基础设施即代码和配置管理工具的经验(我们使用 Terraform 和 Ansible)
- 构建和维护 CI/CD 管道并管理部署的经验
- 精通安全基础知识,包括 Linux 安全加固、网络安全和加密
查看英文原文
About Humble Robotics
Working at Humble Robotics means taking on the biggest change in ground transportation in decades. We're building an autonomous, zero-emissions hauler that dramatically lowers the cost of freight with groundbreaking vision-based AI, designed for today's global logistics network.
We're a fast-moving, close-knit team of AV industry veterans and innovative thinkers. We don't believe culture can be engineered – but when it falls into place, it's a once-in-a-lifetime adventure.
Progress has never felt so present.
Position Overview
We're looking for an ML infrastructure engineer to help design, build, and scale the foundational systems we need to realize our ambitious vision. You'll work on tooling and infrastructure that supports every stage of the ML training flywheel and be an important voice in the technical and organizational decisions that shape our work. From areas spanning vehicle compute to data collection to dataset curation to large-scale model training and deployment, help us build reliable, performant, and secure infrastructure that every team at Humble Robotics can rely on. It's fun here. We are doing cool stuff.
The ideal candidate is a first-principles thinker who is comfortable being a broad generalist. Work on every layer of the stack to help make the software iteration loop as fast and efficient as possible. We're a small team, and your input, experience, and knowledge will play a critical role in shaping every system we build, operate, and depend on to achieve our mission.
Key Responsibilities
- Work on data collection infrastructure that moves sensor data reliably and efficiently from our vehicles into our ML platform
- Develop batch compute pipelines for cataloging, exploring, and curating raw data into high-quality training sets
- Design and scale distributed ML training on our GPU clusters
- Take ownership of performance, observability, efficiency, and security across the full pipeline
- Partner with the ML team to understand their workflows and translate them into reliable infrastructure that accelerates their work
Minimum Qualifications
- Experience building and operating high-availability web services on cloud infrastructure
- Experience with infrastructure-as-code and configuration management tools (we use Terraform and Ansible)
- Experience building and maintaining CI/CD pipelines and managing deployments
- Fluent in security fundamentals including Linux hardening, network security, and cryptographic principles
- Hands-on experience with cluster scheduling systems for running large-scale batch computation
- Comfortable reading, writing, and extending non-trivial code (not just scripting)
- Eligible to work in the United States
Preferred Qualifications
- Hands-on experience managing large, high-performance ML training clusters
- Working knowledge of distributed training frameworks and high-performance networking for ML workloads
- Prior infrastructure experience at an early-stage autonomous vehicle or robotics company
- Comfort operating as an early team member—high ownership, low ego, fast iteration
Compensation
This role is eligible for base salary + benefits + equity compensation. Salary ranges are determined by role, level, and location. Within the range, individual pay is determined by additional factors, including qualifications, skills, experience, and location.
Additional Information
As part of the interview process, we may use Artificial Intelligence (AI) tools to compare your qualifications and experience to the job description. A human reviews all AI output and makes a final hiring decision. Humble Robotics does not rely on the output to make any employment decisions. Some applicants may have a legal right to opt-out of the use of AI as part of our interview process. Contact **legal@humblerobotics.ai** to exercise this right or if you have further questions on the use of AI tools in our hiring process.
Humble Robotics is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, national origin, gender, age, religion, disability, sexual orientation, veteran status, marital status or any other characteristics protected by law. Humble Robotics will consider qualified applicants with arrest and conviction records in a manner consistent with local ordinances.