远程工作雷达

技术支持工程师 – Slurm

Technical Support Engineer – Slurm

开发工程职能支持限定地区(需当地身份)
公司NVIDIA
薪资未公开
工作地点India
地域资格限定地区(需当地身份)
时区要求日间重叠约 6 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 India 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

NVIDIA 三十多年来一直在重新定义计算机图形学、个人电脑游戏和加速计算。这是一种由卓越技术驱动的创新传统,也是充满活力的人才所创造的。如今,我们正利用人工智能的无限潜力,定义计算的新时代——在这个时代,我们的 GPU 成为能够理解世界的计算机、机器人和自动驾驶汽车的大脑。做前所未有的事情需要远见、创新以及世界上最优秀的人才。NVIDIA 的员工沉浸在多元、支持性的环境中,鼓励每个人发挥最佳水平。加入我们,看看你如何对世界产生持久的影响。
Slurm 是许多全球最苛刻的 AI 和高性能计算环境中的关键工作负载管理器。我们正在寻找一位专注于支持 NVIDIA 客户使用 Slurm 的技术支援工程师。你将加入一个专门的 Slurm 专家团队,负责处理复杂的支援案例,并帮助客户运行可靠、高效且高度可扩展的集群。此职位要求有丰富的 Slurm 生产环境经验,并具备在调度器及周围的 Linux、网络、存储、认证、数据库和 GPU 基础设施中诊断问题的能力。

你将负责:

  • 从初步调查到解决,负责客户运行生产环境 AI 和 HPC 集群的 Slurm 支援案例。
  • 诊断涉及 slurmctld、slurmd、slurmdbd、作业调度、节点管理、资源分配、计费、认证和高可用性的复杂问题。
  • 解决 Slurm 的配置和策略功能,包括分区、预留、优先级、公平共享、服务质量、填充、抢占、GRES/TRES、cgroups 和作业约束。
  • 使用日志、诊断数据、配置分析、复现和必要时的源代码调试来调查性能、可靠性和可扩展性问题。
  • 在 Slurm 及其相关依赖项之间隔离问题,包括 Linux、MUNGE、数据库、网络、并行存储、容器、GPU 和集群管理系统。
  • 为客户提供建议,包括 Slurm 配置、升级、操作实践、系统资源管理和生产事故的安全恢复。
  • 与工程团队合作,提供清晰的技术描述、可复现的测试用例和充分支持的缺陷报告。
  • 编写指南、知识库文章和诊断工具
查看英文原文

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for over 30 years. It’s a unique legacy of innovation fueled by great technology—and dynamic people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing—an era in which our GPUs act as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. NVIDIANs immerse themselves in a diverse, supportive environment that encourages everyone to do their best work. Join the team and see how you can make a lasting impact on the world.
Slurm is a critical workload manager for many of the world’s most demanding AI and high-performance computing environments. We are looking for a Technical Support Engineer dedicated to supporting Slurm for NVIDIA’s customers. You will join a dedicated team of Slurm subject-matter guides, owning sophisticated support cases and helping customers run reliable, efficient, and highly scalable clusters. This role requires extensive production experience with Slurm and the ability to diagnose issues across the scheduler and the surrounding Linux, networking, storage, authentication, database, and GPU infrastructure.
What you’ll be doing:

  • Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters.
  • Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability.
  • Solve Slurm configuration and policy features, including partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.
  • Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required.
  • Isolate problems across Slurm and its surrounding dependencies, including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems.
  • Advise customers on Slurm configuration, upgrades, operational practices, managing system resources, and safe recovery from production incidents.
  • Collaborate with engineering teams by producing clear technical descriptions, reproducible test cases, and well-supported defect reports.
  • Develop guides, knowledge-base articles, diagnostic tools, and internal training that strengthen Slurm expertise across the support organization.

What we need to see:

  • BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments including business-critical outage incidents.
  • Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.
  • Capacity to identify sophisticated Slurm incidents independently and guide them to a technically sound resolution.
  • In-depth Linux system-administration and solve experience, including systemd, cgroups, authentication, networking, and database-backed services.
  • Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources.
  • Strong analytical and research skills, showing proficiency in distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems.
  • Excellent written and verbal communication skills, including the ability to turn detailed technical findings into clear explanations and actionable recommendations.

Ways to stand out from the crowd:

  • Experience supporting large-scale Slurm environments containing thousands of nodes or GPUs.
  • Experience diagnosing scheduler performance, job-throughput, controller-load, and database-scaling issues.
  • Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation.
  • Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity.
  • Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位