远程工作雷达

高级解决方案架构师,Infiniband 和网络以太网 - NVIS

Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

AI开发工程限定地区(需当地身份)
公司NVIDIA
薪资未公开
工作地点India
地域资格限定地区(需当地身份)
时区要求日间重叠约 6 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 India 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

NVIDIA 正在寻找高级网络(以太网/IB)解决方案架构师加入其 NVIDIA 基础设施专家团队。全球的学术和商业团体正在使用 NVIDIA 产品来革新深度学习和数据分析,并为数据中心提供动力。加入这个团队,共同打造世界上最大、最快的 AI/HPC 系统!我们正在寻找一位能够在一个以客户为中心的动态团队中工作的人员,需要具备出色的沟通能力。该职位将与客户、合作伙伴和内部团队进行互动,分析、定义并实施大规模网络项目。这些工作的范围包括网络、系统设计和自动化,并作为客户的对外接口!

你将负责:

  • 主要职责包括为新老客户提供 AI/HPC 基础设施构建。
  • 支持大规模 AI 集群的运营和可靠性方面的工作,重点关注大规模性能、实时监控、日志记录和警报。
  • 参与并改进服务的整个生命周期——从构思和设计到部署、运行和优化。
  • 在服务上线后对其进行维护,通过测量和监控可用性、延迟和整体系统健康状况。
  • 向内部团队提供反馈,例如提交 bug、记录解决方法并提出改进建议。

我们需要看到:

  • 计算机科学、电子/计算机工程、物理、数学或相关领域的学士/硕士/博士或同等经验。
  • 至少 5 年以上的网络基础、以太网或 InfiniBand 相关经验。
  • 具有网络交换机/路由器平台如 Cumulus Linux、SONiC、IOS、JunosOS 和 EOS 等的实际操作经验。
  • 精通以太网/InfiniBand/RDMA 核心原理。
  • 熟练掌握端到端 IB/Eth 集群部署、适配器配置和固件维护,并能使用主流 RDMA 测试工具进行专业性能基准测试。
  • 能够独立诊断和排查典型的 IB/Eth 网络异常,包括链路波动、连接失败以及带宽和延迟抖动问题。
  • 掌握实用的 RDMA 网络优化策略,如 QP 调整、MTU 配置和拥塞控制优化。
  • 具有 RDMA 加速业务场景的实际工作经验,包括分布式存储和高性能计算等。
查看英文原文

NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer!
What you'll be doing:

  • Primary responsibilities will include building AI/HPC infrastructure for new and existing customers.
  • Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting.
  • Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
  • Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements.

What we need to see:

  • BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
  • At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World.
  • Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc.
  • Possess solid working knowledge of Ethernet/InfiniBand/RDMA core principles.
  • Be proficient in end-to-end IB/Eth cluster deployment, adapter configuration and firmware maintenance, and able to conduct professional performance benchmarking with mainstream RDMA testing tools.
  • Capable of independently diagnosing and troubleshooting typical IB/Eth network anomalies, including link flapping, connection failure, as well as bandwidth and latency jitter issues.
  • Master practical RDMA network optimization strategies such as QP tuning, MTU configuration and congestion control optimization.
  • Hands-on working experience in RDMA-accelerated business scenarios, including distributed storage and high-performance computing clusters.
  • Extensive experience delivering automated network provisioning solutions using tools like Ansible, Salt, and Python.
  • Ability to develop CI/CD pipelines for network operations.
  • Strong written, verbal, and listening skills in English are essential.

Ways to stand out from the crowd:

  • Familiarity with cloud networks (AWS, GCP, Azure) is a plus.
  • Advanced Linux or Networking Certifications.
  • Experience with High-performance computing architectures. Understanding of how job schedulers(Slurm, PBS) work.
  • luster management technologies knowledge (bonus credit for BCM (Base Command Manager).)
  • Experience with GPU (Graphics Processing Unit) focused hardware/software.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

高级软件开发工程师

NVIDIASwitzerlandFull Time今天
AI开发工程限定地区(需当地身份)日间重叠约 2 小时,需偶尔早起或晚睡

← 返回全部职位