容量运营经理
Capacity Operations Manager
关于 Baseten
Baseten 为全球最具活力的 AI 公司提供关键的推理支持,包括 Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma 和 Writer。通过结合应用 AI 研究、灵活的基础设施和无缝的开发者工具,我们使处于 AI 前沿的公司能够将前沿模型投入生产。我们正在快速成长,并最近完成了 15 亿美元的 F 轮融资 https://www.baseten.co/blog/announcing-our-series-f/,由 Altimeter Capital、Conviction Partners 和 Spark Capital 领投。加入我们,帮助构建工程师用来交付 AI 产品的平台。
职位描述
我们正在寻找一位亲力亲为的运营经理,负责管理 GPU 集群的运营和分析供应端。重点方向包括:GPU 集群生命周期、健康状态、可观测性、利用率监控以及在我们的 neocloud 和裸金属环境中的修复工作。
我们与供应商签订固定计算容量的合同。随着时间推移,GPU 会从健康状态变为不健康状态,该角色的目标是减少这种停机时间,确保任何时候都有尽可能多的 GPU 处于健康状态。
这是一个操作类岗位,不涉及人员管理。你将通过清晰的流程、指标、报告、供应商协调和跨职能协作来推动执行。
职责
核心职责:
- 推动供应商保持尽可能多的 GPU 集群在线并处于健康状态。
- 维护合同容量、分配容量、健康容量和使用容量的实时对账,按供应商和集群划分,最大化健康 GPU 数量。
- 供应商归属的集群健康责任:对每个范围内的供应商负责更换 SLA、平均修复时间(MTTR)和 RMA 周期时间。
- SLA 监控、信用索赔和补救执行:跟踪 SLA 表现是否符合合同条款,提交并跟进信用索赔,并在供应商未达标时推动补救计划。
- 在供应商需要进行维护时推动内部沟通,确保所有 Baseten 利益相关方了解影响可用性的活动。
范围与方法
- 灵活性:此列表涵盖该职位的核心内容,而非全部范围。随着职能发展和新问题出现,你将被要求承担相关的额外工作。
- 主人翁意识:我们需要一个真正以“尽一切努力”为行动原则的人,而不是简历上的空话。如果某项任务超出定义的职责范围,但符合关闭容量缺口的整体目标,那就是你的责任。
查看英文原文
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.
THE ROLE
We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments.
We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment.
This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment..
RESPONSIBILITIES
Core Responsibilities:
- Drive suppliers to keep the maximum amount of the GPU fleet online and healthy.
- Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs.
- Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier.
- SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short.
- Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability.
Scope and Approach
- Flexibility: this list covers the core of the role, not the limit of it. You'll be asked to take on adjacent work as the function evolves and as new gaps surface.
- Ownership mindset: we need someone who treats "whatever it takes" as a genuine operating principle, not a line in a job posting. If something falls outside a defined lane but inside the overall goal of closing the capacity gap, it's yours to pick up.
REQUIREMENTS
- 5 to 10+ years within infrastructure working within the compute lifecycle to maximize functional compute, ideally in a hyperscale, cloud, or large-scale compute environment.
- Direct experience managing GPU, server, or data center hardware supplier relationships. You understand fleet health, RMA processes, and how contracted capacity differs from delivered capacity.
- Highly analytical. You should be comfortable pulling your own data, building your own reports, and generating insights without waiting on someone else to hand you a dashboard.
- Comfortable with ambiguity. Part of the job is figuring out what should exist and building it.
- Strong cross-functional collaboration skills. You'll work closely with finance, infrastructure/engineering, legal, and security on a regular basis.
PREFERRED QUALIFICATIONS:
- Experience at a hyperscaler, neo cloud provider, or AI infrastructure company
- Familiarity with GPU hardware lifecycles (NVIDIA H100/H200/GB200 class systems), power/thermal constraints, and supply chain dynamics for compute.
- Experience running formal supplier corrective actions.
BENEFITS
- Competitive compensation, including meaningful equity
- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).