硬件/软件协同设计工程师 - 3P
Hardware / Software CoDesign Engineer - 3P
关于团队
OpenAI 的硬件部门开发针对先进 AI 工作负载独特需求的芯片和系统级解决方案。该团队负责构建下一代 AI 原生芯片,同时与软件和研究合作伙伴紧密协作,共同设计与 AI 模型高度集成的硬件。除了为 OpenAI 的超算基础设施交付生产级芯片,该团队还创建定制的设计工具和方法,加速创新并实现专门针对 AI 的硬件优化。
关于职位
作为我们硬件优化和协同设计团队的一名工程师,你将与不同供应商共同设计未来硬件,以实现可编程性和性能。你将与我们的内核、编译器和机器学习工程师合作,了解他们与 ML 技术、算法、数值近似、编程表达性以及编译器优化相关的独特需求。你将把这些约束条件传达给各个供应商,以推动和影响未来硬件架构的发展,使其更高效地用于我们的模型训练和推理。如果你对高效地在设备间分布大型语言模型充满热情,处理并优化系统级/机架级的网络瓶颈,并最终定制硬件平台的计算管道和内存层次结构,模拟不同抽象层次的工作负载,并与我们的合作伙伴密切合作,那么这是一个绝佳的机会!
该职位位于加利福尼亚州旧金山。我们采用每周 3 天在办公室工作的混合工作模式,并为新员工提供搬迁协助。
主要职责
- 与我们的硬件供应商共同设计未来硬件,以实现可编程性和性能
- 协助硬件供应商开发最优内核,并在我们的编译器中添加支持
- 为不同硬件配置的关键内核制定性能估计,并推动计算核心和内存层次结构特性的决策
- 在不同抽象层次上构建系统性能模型,并进行分析,以推动扩展、横向扩展和前端网络的决策
- 与机器学习工程师、内核工程师和编译器开发者合作,了解他们对高性能加速器的愿景和需求
- 管理与内部和外部合作伙伴的沟通与协调
- 影响硬件合作伙伴的路线图,使其优化以适应 OpenAI 的需求
查看英文原文
About the Team
OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI.
About the Role
As an Engineer on our hardware optimization and co-design team, you will co-design future hardware from different vendors for programmability and performance. You will work with our kernel, compiler and machine learning engineers to understand their unique needs related to ML techniques, algorithms, numerical approximations, programming expressivity, and compiler optimizations. You will evangelize these constraints with various vendors to develop and influence future hardware architectures towards efficient training and inference on our models. If you are excited about efficiently distributing a large language model across devices, dealing with and optimizing system-wide/rack-wide networking bottlenecks and eventually tailoring the compute pipe and memory hierarchy of the hardware platform, simulating workloads at different abstractions and working closely with our partners, this is the perfect opportunity!
This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.
Key Responsibilities
- Co-design future hardware for programmability and performance with our hardware vendors
- Assist hardware vendors in developing optimal kernels and add support for it in our compiler
- Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory hierarchy features
- Build system performance models at different abstraction levels and carry out analysis to drive decisions on scale up, scale out, front end networking
- Work with machine learning engineers, kernel engineers and compiler developers to understand their vision and needs from high performance accelerators
- Manage communication and coordination with internal and external partners
- Influence the roadmap of hardware partners to optimize them for OpenAI’s workloads.
- Evaluate potential partners’ accelerators and platforms.
- As the scope of the role and team grows, understand and influence roadmaps for hardware partners for our datacenter networks, racks, and buildings.
Qualifications
- 4+ years of industry experience, including experience harnessing compute at scale and optimizing ML platform code to run efficiently on target hardware.
- Strong experience in software/hardware co-design
- Deep understanding of GPU and/or other AI accelerators
- Experience with CUDA, Triton or a related accelerator programming language
- Experience driving Machine Learning accuracy with low precision formats
- Experience with system performance modeling and analysis to optimize ML model deployment
- Strong coding skills in C/C++ and Python
- Are familiar with the fundamentals of deep learning computing and chip architecture/microarchitecture.
- Able to actively collaborate with ML engineers, kernel writers, compiler developers, system engineers, chip architects/microarchitects
Preferred Skills
- PhD in Computer Science and Engineering with a specialization in Computer Architecture, Parallel Computing. Compilers or other Systems
- Strong understanding of LLMs and challenges related to their training and inference
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241.
OpenAI Global Applicant Privacy Policy https://cdn.openai.com/policies/global-employee-and-contractor-privacy-policy.pdf
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.