全球容量经理
Global Capacity Manager
ABOUT BASETEN
Baseten 为全球最具活力的 AI 公司提供关键推理支持,如 Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma 和 Writer。通过结合应用 AI 研究、灵活的基础设施和无缝的开发者工具,我们使处于 AI 前沿的公司能够将前沿模型投入生产。我们正在快速成长,并最近完成了 15 亿美元的 F 轮融资 https://www.baseten.co/blog/announcing-our-series-f/,由 Altimeter Capital、Conviction Partners 和 Spark Capital 领投。加入我们,帮助构建工程师们用来交付 AI 产品的平台。
THE ROLE
作为 Baseten 的全球容量负责人,你将领导公司的“引擎室”,设计、保障并优化支撑客户 AI 工作负载的全球 GPU 集群。你将负责容量管理的全流程,从获取数百万美元的 GPU 集群到构建自动化系统,确保跨多云环境的 99.9% 正常运行时间。
这个职位非常适合希望在高金融资产管理与深度基础设施工程之间架起桥梁的创业者工程师。你将成为全球最先进芯片的集群协调者,确保 Baseten 永远不会出现容量中断,同时保持顶级单位经济性。
明确地说,这是一个高风险的工程角色。你将亲自参与 Kubernetes 协调工作,同时领导专注于下一代硬件(如 NVIDIA 的 Blackwell(B200)架构)的专门小组。
EXAMPLE INITIATIVES
- B200 前沿:为 Baseten 的首批 Blackwell GPU 集群设计基础设施准备和部署策略。
- 全球工作负载协调:构建“多云容量管理”系统,实现客户工作负载在不同区域间无缝迁移,以优化成本和延迟。
- 精准 GPU 分诊:开发基于 Go 的自动化操作符,在一小时内识别、隔离并修复不健康的 H100 节点。
- 智能供应链:与管理层合作,为 Baseten 最大的企业客户争取和预留专用容量。
RESPONSIBILITIES
- 领导专门小组:作为特定 GPU 小组(如 H100 或 B200)的负责人,管理这些资产的全生命周期,包括采购、空域管制和维护。
- 高级协调:执行复杂的工作负载迁移和“粘性”部署排空,确保部署顺利进行。
查看英文原文
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.
THE ROLE
As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments.
This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics.
To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes
orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture.
EXAMPLE INITIATIVES
- The B200 Frontier: Architecting the infrastructure readiness and deployment strategy
for Baseten's first Blackwell GPU clusters.
Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems
to move customer workloads seamlessly across regions to optimize cost and latency.
- Precision GPU Triage: Developing automated Go-based operators to identify, cordon,
and repair unhealthy H100 nodes in under an hour.
The Supply Chain of Intelligence: Partnering with leadership to secure and reserve
dedicated capacity for Baseten’s largest enterprise customers.
RESPONSIBILITIES
- Lead Specialized Pods: Act as the lead for specific GPU pods (e.g., H100 or B200),
managing the full lifecycle of acquisition, air traffic control, and maintenance for those
assets.
- Advanced Orchestration: Execute complex workload migrations and "sticky"
deployment drains, ensuring deployment scheduling rules meet strict regional and
compliance requirements.
- Build for Scalability: Design and implement the "next version" of Baseten’s capacity
management system to handle a 10x increase in GPU volume.
Financial Modeling: Leverage your understanding of unit economics to build ROI
models for GPU spend, ensuring Baseten scales profitably.
- Cross-Team Collaboration: Partner with SRE, Infra, and FDE teams to take discrete
operational tasks off their plate and verify "last mile" follow-through on infrastructure
changes.
- Incident Response: Lead capacity-crunch response by rapidly untainting and re-
coordinating workloads during high-pressure outages.
REQUIREMENTS
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics,
or a related field
- 5+ years of professional work experience in a high-growth environment, preferably at a
hyperscaler (GCP, AWS, Azure) or a specialized GPU provider
- Deep expertise in Kubernetes, including hands-on experience with taints, cordons, node
draining, and custom operators
- Demonstrated experience with Go or Python in a production-level environment
Strong financial literacy and the ability to model complex trade-offs between capacity
reliability and cost
- High tenacity and collaborative mindset
BENEFITS
- Competitive compensation, including meaningful equity
- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).