高级软件工程师,开放Harness工程
Senior Software Engineer, Open Harness Engineering
NVIDIA 在过去 25 多年里一直在改变计算机图形学、个人电脑游戏和加速计算。这是一段独特的创新传统,由伟大的技术和令人惊叹的人才所驱动。如今,我们正在利用人工智能的无限潜力,定义计算的新时代。在这个时代,我们的 GPU 成为能够理解世界的计算机、机器人和自动驾驶汽车的大脑。做以前从未做过的事需要远见、创新以及世界上最优秀的人才。作为一名 NVIDIAN,你将置身于一个多元、支持性的环境中,每个人都能受到激励,发挥最佳水平。加入我们的团队,看看你如何对世界产生持久的影响。
你是否对开源软件、开发工具和代理应用的未来充满热情?我们正在寻找一名高级软件工程师,帮助构建核心库、评估系统和可重用的代理框架功能,使 AI 代理对开发者来说更安全、更快、更容易信任。在这个职位中,你将深入参与开源代理框架、沙盒执行、模型提供者基础设施和代理基准测试,以帮助提升下一代代理的能力。你将把代理评估证据转化为实际改进:诊断失败、向上游贡献代码,并在各种模型中证明质量、可靠性、成本和延迟的提升。加入我们的创意工程团队,共同打造未来自主软件系统的基础技术!
你将负责:
- 定义并演进跨框架、模型和基准环境的开放代理框架评估和改进的技术方案。
- 构建和运营可扩展、可重复的评估系统,涵盖 CI、定时计算、沙盒执行、工件来源、追踪捕获和分析。
- 从追踪、日志、工具结果和基准工件中诊断代理失败,区分框架、评估器、环境、推理和模型的原因。
- 设计控制实验和发布标准,区分真正的框架改进与噪声、基准工件或模型特定的提升。
- 发布聚焦的框架和运行时改进,包括上游开源贡献、检测器、回归测试和可维护的文档。
- 与 Relay、Hermes、模型、基础设施和研究团队合作,将常见的优化需求转化为可重用的运行时、追踪和评估能力。
我们希望看到:
查看英文原文
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
Are you excited about open-source software, developer tools, and the future of agentic applications? We are looking for a Senior Software Engineer to help build the core libraries, evaluation systems, and reusable harness capabilities that make AI agents safer, faster, and easier for developers to trust. In this role, you will work hands-on across open-source agent harnesses, sandboxed execution, model-provider infrastructure, and agent benchmarks to help improve the next generation of agents. You will turn agent-evaluation evidence into real improvements: diagnosing failures, contributing upstream, and proving quality, reliability, cost, and latency gains across models. Join our creative engineering team building foundational technology for the future wave of autonomous software systems!
What you'll be doing:
- Define and evolve the technical approach for evaluating and improving open agent harnesses across frameworks, models, and benchmark environments.
- Build and operate scalable, reproducible evaluation systems spanning CI, scheduled compute, sandboxed execution, artifact provenance, trace capture, and analysis.
- Diagnose agent failures from traces, logs, tool results, and benchmark artifacts, separating harness, evaluator, environment, inference, and model causes.
- Design controlled experiments and promotion criteria that distinguish genuine harness improvements from noise, benchmark artifacts, or model-specific gains.
- Ship focused harness and runtime improvements, including upstream open-source contributions, detectors, regression tests, and maintainable documentation.
- Partner with Relay, Hermes, model, infrastructure, and research teams to turn recurring optimization needs into reusable runtime, trace, and evaluation capabilities.
What we need to see:
- Bachelor’s or Master’s degree in Computer Science, Artificial Intelligence, Applied Math, or a related field, or equivalent experience.
- 8+ years of hands-on software-engineering experience, with demonstrated technical ownership of production systems; and architecture design leadership.
- Expert Python skills and working proficiency in Rust, Go, C++, or TypeScript, with the ability to debug systems across language and process boundaries.
- Experience building or extending LLM agents, coding agents, tool-use loops, model-provider integrations, developer tools, or evaluation infrastructure.
- Solid understanding of asynchronous execution, subprocesses, containers, sandboxing, callbacks, retries, networking, and distributed compute.
- Experience crafting reproducible experiments, benchmark methodology, performance investigations, regression gates, or reliability analysis under nondeterministic conditions.
- A record of shipping and maintaining open-source or developer-facing software with sound testing, API development, documentation, and code reviews.
Ways to stand out from the crowd
- Meaningful contributions to open-source coding agents, harnesses, agent runtimes, evaluation frameworks, developer tools, or observability projects.
- Published research, technical writing, patents, or substantive open-source work in autonomous software engineering, agent evaluation, reliability, tool use, or inference efficiency.
- Experience with SWE-bench, Terminal-Bench, AgentBench, or comparable public agent benchmarks.
- Shipped improvements to context management, tool interfaces, retries, sandboxed repository execution, long-running agent loops, or evaluation environments.
- Experience with trace-analysis and observability systems such as OpenTelemetry, OpenInference, structured event pipelines, exporters, or debugging tools.
Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 18, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.Originally posted on Himalayas