GPT基础设施技术项目经理
GPT Infra Technical Program Manager
关于团队
OpenAI 的工业计算团队负责构建和扩展支持先进 AI 系统的外部基础设施生态系统。我们与超大规模供应商、数据中心提供商、云合作伙伴和战略第三方运营商合作,将合同容量转化为可生产的计算资源。
我们的职责涵盖外部部署的全生命周期:商业对齐、技术准备、网络集成、硬件启用、运营准备和长期扩展策略。
随着 OpenAI 基础设施在全球范围的扩展,我们需要能够将复杂的合作伙伴环境转化为可靠、高速度的训练和推理工作负载容量的领导者。
关于职位
我们正在为 GPT 基础设施团队招聘一名技术项目经理,负责交付直接服务于 OpenAI 模型工作负载的外部计算资源。
在此职位中,您将负责复杂的跨职能项目,将第三方基础设施转化为可大规模使用的 token。您将与工程、容量规划、网络、硬件、财务、产品和外部供应商合作,确保部署的容量转化为实际的生产吞吐量。
该职位位于基础设施执行、系统准备和业务影响的交汇点。成功需要强大的技术理解能力、顶级的项目管理能力和在内部团队和外部合作伙伴之间推动责任落实的能力。
这是一个高可见度的职位,直接影响 OpenAI 全球扩展模型训练和推理的能力。
该职位位于加利福尼亚州旧金山,采用每周 3 天在办公室工作的混合办公模式。提供搬迁协助。
主要职责
- 领导端到端的交付项目,将外部基础设施容量转化为可生产的 token 供应。
- 负责第三方环境中的计算、存储、网络、安全和运营依赖项的准备情况。
- 与内部工程团队和外部合作伙伴建立整合计划,明确里程碑、负责人、风险和关键路径。
- 推动新合作伙伴区域、集群和容量扩展的上线执行。
- 建立衡量已部署容量与可用 token 输出的操作机制。
- 识别阻碍 token 生成的瓶颈(如网络限制、硬件准备、软件启用、合作伙伴延迟等)并推动解决。
查看英文原文
About the Team
OpenAI’s Industrial Compute team is responsible for building and scaling the external infrastructure ecosystem that powers advanced AI systems. We work across hyperscalers, colocation providers, cloud partners, and strategic third-party operators to turn contracted capacity into production-ready compute.
Our scope spans the full lifecycle of external deployments: commercial alignment, technical readiness, network integration, hardware enablement, operational readiness, and long-range scaling strategy.
As OpenAI’s infrastructure footprint expands globally, we need leaders who can convert complex partner environments into reliable, high-velocity capacity for training and inference workloads.
About the Role
We are seeking a Technical Program Manager for our GPT Infrastructure teams, to lead delivery of external compute capacity that directly serves OpenAI model workloads.
In this role, you will own complex cross-functional programs that transform third-party infrastructure into usable tokens at scale. You will partner across engineering, capacity planning, networking, hardware, finance, product, and external providers to ensure that deployed capacity translates into real production throughput.
This role sits at the intersection of infrastructure execution, systems readiness, and business impact. Success requires strong technical fluency, elite program management, and the ability to drive accountability across internal teams and external partners.
This is a high-visibility role with direct impact on OpenAI’s ability to scale model training and inference globally.
This role is based in San Francisco, CA, with a hybrid work model of 3 days in office per week. Relocation assistance is available.
Key Responsibilities
- Lead end-to-end delivery programs that convert external infrastructure capacity into production-ready token supply.
- Own readiness across compute, storage, networking, security, and operational dependencies for third-party environments.
- Build integrated plans across internal engineering teams and external partners with clear milestones, owners, risks, and critical paths.
- Drive launch execution for new partner regions, clusters, and capacity expansions.
- Create operating mechanisms that measure deployed capacity versus usable token output.
- Identify bottlenecks preventing token generation (network constraints, hardware readiness, software enablement, partner delays, etc.) and drive resolution.
- Coordinate with capacity planning and finance teams to prioritize the highest ROI capacity opportunities.
- Establish executive-level reporting on delivery status, risks, and token ramp forecasts.
- Improve repeatability of partner onboarding, technical integration, and scaling motions.
- Manage escalations across internal and external stakeholders during high-severity delivery issues.
- Translate ambiguous infrastructure constraints into clear execution plans.
- Help define the long-term operating model for Token-as-a-Service across Stargate and 3P ecosystems.
Qualifications
- 8+ years of Technical Program Management, Engineering Program Management, or Infrastructure Delivery experience.
- Experience leading large-scale technical programs involving cloud, data center, networking, hardware, or distributed systems.
- Strong understanding of compute infrastructure, clusters, networking, storage, and production systems.
- Proven ability to drive cross-functional execution across engineering, operations, finance, and external vendors.
- Experience managing executive stakeholders and communicating complex tradeoffs clearly.
- Strong analytical skills with ability to reason about utilization, throughput, capacity, and operational metrics.
- Comfortable operating in ambiguous, fast-scaling environments.
- Strong written and verbal communication skills.
- High ownership mentality with bias toward action.
- Experience working with external providers, strategic partners, or hyperscalers is highly preferred.
Preferred Skills
- Experience with GPU clusters, AI infrastructure, or large-scale model serving environments.
- Familiarity with token economics, inference capacity planning, or workload scheduling.
- Experience scaling global infrastructure through third-party providers.
- Background in systems engineering, networking, or hardware deployment programs.
- Experience building new operational models in high-growth environments.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241.
OpenAI Global Applicant Privacy Policy https://cdn.openai.com/policies/global-employee-and-contractor-privacy-policy.pdf
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.