高级站点可靠性工程师,DGX云
Senior Site Reliability Engineer, DGX Cloud
NVIDIA正在推动人工智能和高性能计算的发展。DGX Cloud旨在为主要云服务提供商提供一个完全托管的人工智能平台,利用高性能的NVIDIA基础设施优化人工智能工作负载。作为高级站点可靠性工程师,与NVIDIA的DGX Cloud团队合作,为全球的人工智能研究人员和企业客户维护高性能的DGX Cloud集群。
这个机会之所以出色,是因为你将处于技术的最前沿,与创新的人工智能和云计算解决方案一起工作。你将有机会加入一个致力于突破创新边界并完美实施雄心勃勃项目的顶级团队!
你将负责:
- 构建、实施并支持大规模Kubernetes集群的运营和可靠性方面,重点关注性能、实时监控、日志记录和警报。
- 定义SLO/SLI,监控错误允许量,并优化报告流程。
- 在服务发布前通过系统创建咨询、开发软件工具、平台和框架、容量管理以及发布评审提供支持。
- 服务上线后通过测量和监督可用性、延迟和整体系统健康状况进行维护。
- 在AWS、GCP、Azure、OCI和私有云上操作和优化GPU工作负载。
- 通过自动化等机制可持续扩展系统,并通过推动改进可靠性和速度的变更来演进系统。
- 领导高严重性事件的故障排查和根本原因分析。
- 实践平衡的事件响应和无责复盘。
- 参与值班轮班以支持生产服务。
我们希望看到:
- 计算机科学或相关技术领域的学士学位,或同等经验。
- 8年以上运行生产服务的经验。
- 精通Kubernetes管理、容器化和微服务架构。
- 使用基础设施自动化工具(如Terraform、Ansible、Chef、Puppet)的经验。
- 精通至少一种高级编程语言(如Python、Go)。
- 对Linux操作系统、网络基础(TCP/IP)和云安全标准有深入理解。
- 熟悉SRE原则,如SLO、SLI、错误预算和事件管理。
- 有构建和运行全面可观测性堆栈(监控、日志、追踪)的经验,使用工具如OpenTelemetry、Prometheus、Grafana等。
查看英文原文
NVIDIA is driving AI and high-performance computing forward. DGX Cloud aims to deliver a fully managed AI platform on major cloud providers, optimizing AI workloads using high-performance NVIDIA infrastructure. Work with NVIDIA's DGX Cloud team as a Senior Site Reliability Engineer to maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide.
What makes this opportunity outstanding is that you will be at the forefront of technology, working with innovative AI and cloud computing solutions. You will have the chance to contribute to a world-class team that is determined to push the boundaries of innovation and flawlessly implement ambitious projects!
What you’ll be doing:
- Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting.
- Define SLOs/SLIs, monitor error allowances, and streamline reporting.
- Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews.
- Maintain services once they are live by measuring and supervising availability, latency, and overall system health.
- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
- Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
- Lead triage and root-cause analysis of high-severity incidents.
- Practice balanced incident response and blameless postmortems.
- Participate in on-call rotation to support production services.
What we need to see:
- BS in Computer Science or related technical field, or equivalent experience.
- 8+ years of experience operating production services.
- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
- Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet).
- Proficiency in at least one high-level programming language (e.g., Python, Go).
- In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards.
- Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.
- Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
Ways to stand out from the crowd:
- Operating GPU-accelerated clusters with KubeVirt in production.
- Applying generative-AI techniques to reduce operational toil.
- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.
- Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 19, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.Originally posted on Himalayas