技术项目经理(硬件)
Technical Project Manager (Hardware)
关于Nebius:
Nebius正在引领全球AI经济的云基础设施新时代。我们正在构建一个全栈AI云平台,支持开发者和企业从数据和模型训练到生产部署的全流程,无需承担自建大型内部AI/ML基础设施的成本和复杂性。
由工程师打造,面向工程师。从大规模GPU编排到推理优化,我们在计算、存储、网络和应用AI领域掌握着关键难题。
在纳斯达克上市(NBIS),总部位于阿姆斯特丹,我们在欧洲、英国、北美和以色列设有研发中心,拥有全球业务覆盖。我们的团队超过1,500人,包括数百名在硬件、软件和AI研发方面具有深厚专业知识的工程师。
职位描述
我们正在寻找一名技术项目经理,负责大规模AI数据中心基础设施的服务器和机架级别的生产项目。该职位需要与容量规划、硬件架构、采购、ODM制造、物流和数据中心部署团队进行协作。你将确保正确的GPU、计算和存储基础设施按时订购、生产、交付并部署。该职位需要扎实的服务器硬件知识、结构化的项目管理能力、供应商协调能力,以及对相关工作流的高度责任感。
你的职责包括:
- 与容量规划和数据中心架构团队对齐服务器和机架需求。将项目需求转化为采购和ODM生产的硬件输入。
- 分解所需的关键部件,包括CPU、内存、硬盘、网卡和其他关键组件。验证采购的组件是否符合所需的服务器配置。
- 与ODM合作伙伴下订单并跟踪生产进度。协调物料清单(BOM)准备、组件可用性和生产时间表。
- 管理多个项目和供应商的产品生产组合。跟踪风险、障碍、依赖关系和恢复计划。
- 早期上报问题,包括BOM问题、缺货、质量风险、工厂延误、付款延迟、出口许可证问题和部署时间表变更。必要时前往工厂以支持质量控制、生产执行和时间表恢复。
- 跟踪相关流程,如采购订单(PO)、发票、出口许可证、付款、物流预订和所需文件。
- 向相关方定期汇报生产状态、风险更新和决策点。
- 推动问题解决
查看英文原文
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
We are looking for a Technical Project Manager to own server and rack-level production projects for large-scale AI datacenter infrastructure. This role connects capacity planning, hardware architecture, procurement, ODM manufacturing, logistics, and datacenter deployment teams. You will make sure the right GPU, compute, and storage infrastructure is ordered, produced, delivered, and deployed on time. The role requires strong server hardware knowledge, structured project management, vendor coordination, and high ownership across adjacent work-streams.
Your responsibilities will include:
- Align server and rack requirements with capacity planning and datacenter architecture teams. Translate project demand into hardware inputs for procurement and ODM production.
- Decompose required key parts, including CPUs, memory, drives, NICs, and other critical components. Validate that procured components are compatible with required server configurations.
- Place and track production orders with ODM partners. Coordinate BOM readiness, component availability, and production timelines.
- Manage a portfolio of production projects across multiple projects and vendors. Track risks, blockers, dependencies, and recovery plans.
- Escalate issues early, including BOM problems, shortages, quality risks, factory delays, payment delays, export license issues, and deployment schedule changes. Visit factories when needed to support quality control, production execution, and timeline recovery.
- Track adjacent processes such as POs, invoices, export licenses, payments, logistics bookings, and required documentation.
- Provide regular production status, risk updates, and decision points to stakeholders.
- Drive issue resolution across planning, procurement, production, testing, logistics, delivery, and deployment.
- Continious improvements after each production cycle. Improve ODM requirements, supplier expectations, internal processes, reporting, and tooling. Gather deployment and operations feedback and convert it into future production improvements.
We expect you to have:
- Experience managing complex hardware production, manufacturing, supply chain, infrastructure, or datacenter delivery projects.
- Strong understanding of modern server architectures, especially high-performance GPU-based environments.
- Knowledge of server components and ability to configure hardware for different service needs.
- Ability to quickly adapt to new server platforms, components, and technologies.
- Experience working with ODMs, OEMs, factories, hardware vendors, or contract manufacturers.
- Experience with BOMs, component planning, production timelines, and factory execution.
- Strong project management skills, including Gantt charts, milestones, dependencies, and risk tracking.
- Excellent problem-solving skills and ability to work under tight deadlines.
- Strong communication, negotiation, and vendor management skills.
- High ownership, structured thinking, and ability to drive work-streams outside the narrow TPM scope.
It will be an added bonus if you have:
- Basic understanding of GPU infrastructure including diagnostics/troubleshooting of GPU.
- Ability to collect system logs and hardware diagnostics, with a general understanding of BIOS/BMC functionality and their interaction with the operating system and hardware components.
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
If you need accommodations during the application process, please let us know.