软件工程师 - 培训基础设施
Software Engineer - Training Infrastructure
ABOUT BASETEN
Baseten 为全球最具活力的 AI 公司提供关键任务的推理支持,例如 Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma 和 Writer。通过结合应用 AI 研究、灵活的基础架构和无缝的开发者工具,我们使处于 AI 前沿的公司能够将前沿模型投入生产。我们正在快速成长,并最近完成了 1.5 亿美元的 F 轮融资 https://www.baseten.co/blog/announcing-our-series-f/,由 Altimeter Capital、Conviction Partners 和 Spark Capital 领投。加入我们,帮助构建工程师们用来交付 AI 产品的平台。
THE ROLE
作为 Training Infrastructure 团队的软件工程师,你将设计并主导我们训练平台的开发,支持顶级的研究工程师和模型开发者。你将为支撑开发者部署、扩展和监控工作负载的基础设施做出关键的技术决策。你将负责训练堆栈中技术系统的调度、存储、网络、可靠性及可观测性
EXAMPLE INITIATIVES
看看我们迄今为止打造的产品:
- 当前产品的概述 https://www.baseten.co/blog/baseten-training-is-ga/#training-is-now-ga
- 训练文档概述 https://docs.baseten.co/training/overview
- 训练产品的故事 https://www.baseten.co/blog/a-q-a-from-inference-to-training-the-inside-story-of-baseten-s-newest-product/
- 我们所做的研究 https://www.baseten.co/resources/research/
RESPONSIBILITIES
- 设计并构建可扩展的基础设施系统,用于我们的 ML 训练平台(如调度、存储和网络)
- 与开发者和研究工程师紧密合作,将复杂的训练需求转化为技术解决方案
- 设计并构建全球训练调度器
- 设计并构建强化学习系统和持续学习流水线
- 推动长期改进,以提高系统的可靠性和发展速度
- 与 SRE 和 Capacity 团队紧密合作,实现最先进的训练基础设施
- 在性能与系统可靠性之间做出关键的架构决策
- 主导技术讨论,指导初级工程师掌握基础设施的最佳实践
- 参与长期技术战略和基础设施路线图的制定
REQUIREMENTS
- 计算机科学或相关领域的学士学位或更高
查看英文原文
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.
THE ROLE
As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack
EXAMPLE INITIATIVES
Take a look at what we’ve built so far:
- Overview of the product so far https://www.baseten.co/blog/baseten-training-is-ga/#training-is-now-ga
- Training docs overview https://docs.baseten.co/training/overview
- Story of the Training product https://www.baseten.co/blog/a-q-a-from-inference-to-training-the-inside-story-of-baseten-s-newest-product/
- Research we've done https://www.baseten.co/resources/research/
RESPONSIBILITIES
- Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking)
- Partner closely with developers and research engineers to translate complex training requirements into technical solutions
- Design and architect a global training scheduler
- Design and architect reinforcement learning systems and continuous learning pipelines
- Drive long-term improvements to improve reliability of systems and velocity of development
- Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure
- Make critical architectural decisions balancing performance with system reliability
- Lead technical discussions and mentor junior engineers on infrastructure best practices
- Contribute to long-term technical strategy and infrastructure roadmap
REQUIREMENTS
- Bachelor’s degree or high in Computer Science or related field
- Proficiency in Go, with
- Deep expertise with Kubernetes in production environments
- Advanced understanding of distributed systems concepts and performance tuning
- Proven experience designing observability systems
- Experience with ML/AI workloads and MLOps platforms
NICE TO HAVE
- Experience with distributed storage systems
- Python experience a plus
- Extensive experience with major cloud providers (AWS, GCP) and neo-cloud providers (Crusoe, DigitalOcean, Nebius)
- Experience with workload orchestration platforms like Temporal or Airflow
- Familiarity or experience with the open source training stack and frameworks (NCCL, PyTorch, Megatron, NemoRL, VeRL, Axolotl, HF Trainer) and distributed training techniques (FSDP, DeepSpeed).
- Experience developing AI products, tooling, or agents
BENEFITS
- Competitive compensation, including meaningful equity
- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).