高级系统软件集成与调试工程师
Senior System Software Integration and Debug Engineer
NVIDIA在计算机图形学、个人电脑游戏和高性能计算领域的影响已经持续了25年以上。这是一种由卓越技术和杰出人才驱动的创新传统。目前,我们正在利用人工智能的无限可能,打造计算的新时代。在这个时代,我们的GPU将成为计算机、机器人和自动驾驶汽车感知世界的核心智能。创造前所未有的事物需要创造力、创新精神以及全球顶尖的人才。作为NVIDIA的一员,你将加入一个多元化、积极向上的工作环境,鼓励每个人发挥最高水平。加入团队,发现你如何对世界产生持久影响。
NVIDIA正在寻找高度积极、富有创造力的工程师加入平台软件团队。你将与一支勤奋的软件工程师团队合作,涉及SOC、系统和各类技术垂直领域的各个方面。作为一名勤奋且热爱自己专业的工程师,你将负责设计SOC驱动程序、BSP、复杂的CI/CD系统,并与关键合作伙伴和OEM客户协作。你需要在快节奏和敏捷的环境中表现出色。
你将负责:
- 负责跨固件、BSP、驱动程序、GPU软件、服务、测试和发布系统的端到端软件集成策略。
- 主导系统级启动失败、挂起、崩溃、意外重启、电源转换失败、性能退化、更新失败以及固件/驱动程序交互的调试。
- 使用崩溃转储、事件日志、跟踪信息、UART输出、固件日志、遥测数据和硬件诊断来隔离系统各层的故障。
- 构建自动化故障分类系统,区分产品缺陷与基础设施、配置、设备、测试框架和不稳定测试的故障。
- 推动逃逸缺陷、集成失败和重复基础设施问题的根本原因分析和纠正措施。
- 将OEM和合作伙伴的软件集成到NVIDIA软件栈中,确保生产平台的合规性、稳定性以及功耗和性能基准。
- 与多职能团队紧密合作,优先处理所有BSP组件(固件、驱动程序、内核和硬件层)中的问题。
- 与架构、芯片、固件和操作系统工程团队协作,实现新功能并确保跨编译的无缝性。
查看英文原文
NVIDIA’s impact on computer graphics, personal computer gaming, and high-performance computing has been evolving for over 25 years. It’s a distinctive heritage of innovation that’s motivated by great technology—and remarkable people. Currently, we’re harnessing the boundless possibilities of AI to build the next era of computing. An era where our GPU functions as the intelligence behind computers, robots, and autonomous vehicles that perceive the world. Creating what’s never existed calls for creativity, inventiveness, and the world’s leading talent. As an NVIDIAN, you’ll be part of a varied, encouraging workplace that encourages everyone to perform at their highest level. Join the team and discover how you can build a lasting impact on the world.
NVIDIA is searching for highly motivated, creative engineers to join the Platform Software team. You will work with a team of hardworking software engineers across all aspects of SOC, systems, and technology verticals. As someone who is hardworking and passionate about their craft, you will design key aspects of our SOC drivers, BSP, sophisticated CI/CD system, as well as collaborating with key partners and OEM customers. You should demonstrate the ability to excel in an environment with fast pace and agility.
What you'll be doing:
- Own the end-to-end software integration strategy for products across firmware, BSP, drivers, GPU software, services, test and release systems.
- Lead system-level debugging of boot failures, hangs, crashes, unexpected resets, power-transition failures, performance regressions, update failures, and firmware/driver interactions.
- Use crash dumps, event logs, traces, UART output, firmware logs, telemetry, and hardware diagnostics to isolate failures across system layers.
- Build automated failure classification that distinguishes product defects from infrastructure, configuration, device, test-harness, and flaky-test failures.
- Drive root-cause analysis and corrective actions for escaped defects, integration failures, and recurring infrastructure problems.
- Integrate OEM and partner software into the NVIDIA software stack to ensure compliance, stability, and power/performance benchmarks for production platforms.
- Work closely with multi-functional teams to prioritize issues across all BSP components—firmware, drivers, kernel, and hardware layers.
- Collaborate with architecture, silicon, firmware, and OS engineering teams to enable new features and ensure flawless cross-component integration.
What we need to see:
- BS or MS degree in Computer Engineering, Computer Science, or equivalent experience, with 12+ years of relevant software development experience.
- Strong understanding of ARM microarchitecture and exception levels, with an emphasis on integration, triage, and debugging.
- Understanding of SoC architecture spanning Boot, Security, Power, and OS bring-up. Good understanding of ACPI and Device Tree concepts.
- Proficiency in C/C++ and Python for automation and validation tooling.
- Solid understanding of Kernel and Hypervisor internals on both Windows and Linux systems.
- Experience integrating drivers/firmware and debugging kernel components, with a specific focus on Windows. Background in solving problems within large, complex systems deployed at scale.
Ways to stand out from the crowd:
- Background and strength with sophisticated system-level debugging is invaluable
- Experience working on system-level reliability and resiliency features.
- Experience products integration and deliverables E2E
- Familiarity with system-level security concepts
- Experience with embedded system SW concepts.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 21, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.Originally posted on Himalayas