远程工作雷达

高级解决方案架构师,AI计算工程师 - NVIS

Senior Solution Architect, AI Compute Engineer - NVIS

AI开发工程限定地区(需当地身份)
公司NVIDIA
薪资未公开
工作地点Australia
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 Australia 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

NVIDIA 在过去 25 多年里一直在改变计算机图形、PC 游戏和加速计算。这是一段独特的创新历史,由伟大的技术和令人惊叹的人才所驱动。如今,我们正在利用人工智能的无限潜力,定义计算的新纪元。在这个纪元中,我们的 GPU 成为能够理解世界的计算机、机器人和自动驾驶汽车的大脑。做以前从未做过的事情需要远见、创新和世界上最优秀的人才。作为 NVIDIA 的一员,你将置身于一个多元、支持性的环境中,每个人都会被激励去做出最好的工作。加入团队,看看你如何对世界产生持久的影响。
NVIDIA 正在寻找一名高级 AI/HPC 工程师加入其基础设施专家团队。全球的学术和商业团体都在使用 NVIDIA 的产品来革新深度学习和数据分析,并为数据中心提供动力。加入构建世界上最大、最快的 AI/HPC 系统的团队!NVIDIA 正在寻找一位能够在动态、以客户为中心的团队中工作的人员,该职位需要出色的沟通能力。此职位将与客户、合作伙伴和内部团队互动,分析、定义并实施大规模的 AI/HPC 项目。这些工作的范围包括网络、系统设计和自动化,并且是客户的对接人。
你将负责的工作:

  • 主要职责包括在基于 Linux 的环境中为新老客户提供 AI/HPC 基础设施的部署、管理和维护。
  • 在规划会议到实施过程中,作为客户领域的专家。
  • 提供相关的交接文档,并进行知识转移,以支持客户开始部署世界上最复杂的系统之一!
  • 向内部团队提供反馈,例如提交 bug、记录解决方法并提出改进建议。

我们希望看到:

  • 计算机科学、电子/计算机工程、物理、数学或相关领域的学士/硕士/博士或同等经验。
  • 5 年以上提供深入支持和部署服务的经验,解决软硬件产品的相关问题。
  • 熟悉 Linux 系统管理、进程管理、包管理、任务调度、内核管理、启动流程/故障排查、性能报告/优化/日志记录、网络路由/高级功能。
查看英文原文

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
NVIDIA is looking for a Senior AI/HPC Engineer to join its infrastructure Specialist team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! NVIDIA is looking for someone with the ability to work on a dynamic, customer-focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners, and internal teams to analyse, define, and implement large-scale AI/HPC projects. The scope of these efforts includes a combination of Networking, System Design, and Automation, and being the face to the customer.
What you will be doing:

  • Primary responsibilities will include deploying, managing and maintaining AI/HPC infrastructure in Linux-based environments for new and existing customers.
  • Be the domain expert with customers during planning calls through implementation.
  • Handover-related documentation and perform knowledge transfers required to support customers as they begin rolling out some of the most sophisticated systems in the world!
  • Provide feedback into internal teams such as opening bugs, documenting workarounds, and suggesting improvements.

What we need to see:

  • BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
  • 5+ years providing in-depth support and deployment services, solving problems for hardware and software products.
  • Knowledge and experience with Linux System Administration, process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging, network routing/advanced networking (tuning and monitoring).
  • Cluster management technologies.
  • Scripting proficiency.
  • Good interpersonal skills with the ability to maintain and deliver resolutions for customer-blocking issues as they arise. Excellent verbal and written English skills.
  • Strong organizational skills and ability to prioritize/multi-task easily with limited supervision.
  • Industry-standard Linux certifications.
  • Experience with Schedulers such as SLURM, LSF, UGE, etc.

Ways to stand out from crowd:

  • Demonstrated hands-on experience with MPI (e.g., OpenMPI, MPICH), proficient in distributed communication programming and cluster debugging.
  • In-depth understanding of NCCL principles and applications, with expertise in collective communication optimization for NVIDIA GPU clusters.
  • Experience in deploying and optimizing high-speed networks (InfiniBand/Ethernet), with a clear understanding of how network architecture impacts GPU cluster performance
  • Familiarity with automation tools (Ansible, Salt, Puppet, etc.), capable of implementing batch configuration and operational automation for GPU clusters and LLM deployment environments.
  • knowledge and hands-on experience with Kubernetes, including container orchestration for AI/ML workloads, resource scheduling, scaling, and integration with HPC environments.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. Are you a creative and autonomous manager who loves a challenge? Are you ready to become the staff member you always wanted to be? Come and be part of the best SW design team in the industry!
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

开发者技术经理

NVIDIACanada, United States$224,000 - $431,250/年Full Time今天
开发工程限定地区(需当地身份)

高级空中平台软件工程师

NVIDIASwedenFull Time今天
开发工程职能支持限定地区(需当地身份)日间重叠约 2 小时,需偶尔早起或晚睡

← 返回全部职位