技术项目经理,AI工厂基础设施
Technical Program Manager, AI Factory Infrastructure
NVIDIA的AI工厂基础设施团队开发全球参考设计,降低下一代计算和网络产品所需技术的风险,并构建大规模AI基础设施以验证解决方案。作为该团队的TPM,你将负责领导高度跨职能的团队,包括硬件、软件和设施基础设施团队,以交付解决方案和基础设施。这个职位提供了真正独特的结合:既参与未来解决方案和产品的开发,又实现大规模基础设施的构建。
你将负责的工作:
- 与产品负责人和技术负责人合作,识别并收集下一代AI工厂的需求。
- 建立、监督并完成长期项目,包括时间表、资源分配和检查点。
- 与数据中心和硬件团队合作,找到解决难题的创新方案,并共同开发解决方案和缓解策略。
- 与关键内部合作伙伴就容量需求进行规划,结合工程路线图和数据中心扩展。
- 负责100MW以上AI工厂数据中心部署的端到端交付,从建设准备到调试、启动和移交给运营。
- 协调总承包商、托管服务提供商、公用事业公司和OEM,协调电力/机械范围、长周期设备和现场物流,以实现大规模AI集群部署。
- 在电源、液冷、网络和平台/硬件团队之间推动集成准备审查和验收标准,确保AI工厂应用的性能和可靠性目标达成。
- 为政府资助和计划制定项目计划。
- 将项目需求转化为设计基础文档。
- 汇聚团队成员,促进协作交付方式,同时对团队成员的任务和时间表保持问责。
我们希望看到:
- 出色的长期规划和执行能力,以推进数据中心生命周期规划,包括大规模AI工厂的建设和扩展。
- 管理高密度AI基础设施数据中心端到端部署的经验,包括调试、准备审查、启动和运营交接。
- 在协调托管服务提供商、总承包商、公用事业公司和OEM方面有成功经验,以交付复杂的电力和机械范围(例如在100MW以上园区规模)。
- 强大的技术和项目领导能力
查看英文原文
NVIDIA's AI Factory Infrastructure team develops global reference designs, de-risks technologies needed for the next generation of compute and network products and builds AI infrastructure at scale to validate solutions at scale. As a TPM on the team, you’ll be responsible for leading teams that are highly multi-functional, including hardware, software and facility infrastructure teams, to deliver solutions and infrastructure. This role offers a truly unique blend of developing future solutions and products with building infrastructure at scale.
What You Will Be Doing:
- Collaborate with product owners and technical leads to identify and collect requirements for next-generation AI Factories.
- Build, supervise, and complete long-term programs including schedules, resourcing, and checkpoints.
- Work with data center and hardware teams to find creative solutions to hard problems, and co-develop solutions and mitigation strategies.
- Lead planning with key internal partners on capacity demands with engineering roadmaps and data center expansions.
- Own end-to-end delivery of 100MW+ AI factory data center deployments, from construction readiness through commissioning, turn-up, and handoff to operations.
- Coordinate general contractors, colocation providers, utilities, and OEMs to align electrical/mechanical scope, long-lead equipment, and site logistics for large-scale AI cluster deployments.
- Drive integrated readiness reviews and acceptance criteria across power, liquid cooling, networking, and platform/hardware teams to ensure performance and reliability targets are met for AI factory applications.
- Develop program plans for government grants and initiatives.
- Translate program requirements into Basis Of Design documents
- Bring together team members and foster a collaborative approach to delivery while holding team members accountable to action items and timelines.
What We Need To See:
- Outstanding long-term planning and execution skills to carry our data center lifecycle planning, including large-scale AI factory buildouts and expansions.
- Experience managing end-to-end data center deployments for high-density AI infrastructure, including commissioning, readiness reviews, turn-up, and operational handoff.
- Demonstrated ability to coordinate across colocation providers, general contractors, utilities, and OEMs to deliver complex electrical and mechanical scope (e.g., at 100MW+ campus scale).
- Strong technical and program leadership across power delivery, liquid cooling, networking, and compute/platform teams to define acceptance criteria and ensure performance and reliability targets are met.
- 12+ years of experience providing program and project management leadership for data center projects covering construction of mechanical, electrical, and plumbing with large-scale server, storage, and network deployments.
- BS or MS degree in Engineering (or equivalent experience).
Ways To Stand Out From the Crowd:
- In-depth knowledge of infrastructure (hardware and software) data center facilities infrastructure (electrical and mechanical) technologies.
- Familiarity with NVIDIA’s AI compute technology stack and ability to translate platform requirements into data center infrastructure designs (power delivery, liquid cooling, space, and network topology) at scale.
- Experience with colocation data center environments
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD for Level 5, and 240,000 USD - 379,500 USD for Level 6.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 24, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.Originally posted on Himalayas