软件工程师 - Baseten 推理堆栈
Software Engineer - Baseten Inference Stack
关于Baseten
Baseten为全球最具活力的AI公司提供关键任务的推理服务,包括Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma和Writer。通过结合应用AI研究、灵活的基础设施和无缝的开发者工具,我们使处于AI前沿的公司能够将最前沿的模型投入生产。我们正在快速成长,并最近完成了15亿美元的F轮融资,由Altimeter Capital、Conviction Partners和Spark Capital领投。加入我们,帮助构建工程师们用来交付AI产品的平台。
职位描述
Baseten的推理栈团队构建了支撑我们平台上大规模LLM推理的分布式运行时。我们在分布式系统、模型性能、基础设施和开发者体验的交汇点上工作。我们使客户能够以行业领先的性能、可扩展性、可靠性和易用性部署和运营最先进的LLM模型。
作为推理栈团队的软件工程师,你将跨整个技术栈工作——从客户用于部署模型的开发者体验,到用于工具调用和推理等功能的库,一直到我们用于在Kubernetes中编排部署和高效路由流量的系统。
这是一个适合喜欢在生产环境中负责系统、解决复杂集成问题,并为用户简化和提升复杂基础设施可靠性的工程师的理想职位。
示例项目
博客文章
https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/
https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/
https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid
职责
- 开发用于部署和管理大规模分布式LLM推理的基础架构和编排系统
- 跨技术栈工作,从面向客户的特性到底层基础设施组件
- 构建与路由、自动扩展、调度、可观测性和运行时管理相关的平台功能
- 提升我们推理栈的可靠性、可扩展性和可用性
- 与模型性能工程师紧密合作,将新的推理优化广泛地提供给客户并易于配置
- 帮助制定测试方面的最佳实践,以及重新设计和改进现有系统
查看英文原文
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.
THE ROLE
Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use.
As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently.
This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users.
EXAMPLE INITIATIVES
Blog Posts
https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/
https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/
https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid
RESPONSIBILITIES
- Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference
- Work across the stack, from customer-facing features to low-level infrastructure components
- Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management
- Improve the reliability, scalability, and usability of our inference stack
- Collaborate closely with Model Performance engineers to make new inference optimizations broadly available to customers and easy to configure
- Help define best practices around testing, release automation, benchmarking, and operational excellence
- Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads
- Make thoughtful engineering tradeoffs balancing performance, reliability, operational simplicity, and developer experience
- Own projects end-to-end: from architecture and implementation through deployment, monitoring, and iteration based on customer feedback
REQUIREMENTS
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field
- Strong background in distributed systems, backend infrastructure, or platform engineering
- Experience building and operating production systems where reliability, latency, and scale are first-class concerns
- Strong sense of developer experience: you think about how systems are used, not just how they work
- Motivated and willing to learn new languages, frameworks, and systems as needed
- Ability to debug complex systems across multiple layers of the stack
- Genuine interest in inference engineering. You don’t need to have hands on experience but are willing to learn
- Excellent communication and collaboration skills
NICE TO HAVE
- Experience with Kubernetes, including concepts like operators and custom resources
- Prior work on Dynamo, vLLM, SGLang, TensorRT-LLM, or similar inference frameworks
- Experience with distributed scheduling, autoscaling, or service orchestration
- Experience operating GPU workloads in production
- Familiarity with observability tooling, CI/CD systems, or release automation
- Experience contributing to open-source infrastructure or ML systems
BENEFITS
- Competitive compensation, including meaningful equity
- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).