高级系统软件工程师,软件定义网络
Senior System Software Engineer, Software-Defined Networking
在NVIDIA Gaming Network(NGN)基础设施/存储团队中,NGN软件定义存储(SDS)团队专注于NGN/GeForce NOW的存储和基础设施软件,工作内容涵盖NGN 2.x Kubernetes、块存储、性能/扩展性、安全性、租户存储、租户游戏座位以及相关的中心NGN服务。
NVIDIA正在寻找一名高级系统软件工程师(SDN),以推动全球GPU云基础设施的软件定义网络,该基础设施支持AI工作负载、云游戏、内容分发和其他加速服务。使用Open vSwitch、OVN、Kubernetes、C和Go构建和运营安全、可编程的云网络。与高级工程师、NVIDIA架构师和合作伙伴团队协作,主导架构设计、实现、验证、问题响应和生产改进。
你将负责:
- 与工程师和架构师合作,使用Open vSwitch、OVN、OpenFlow和覆盖网络技术,定义、评审并演进多租户控制平面和数据平面架构。
- 通过设计和代码评审、实现指导、生产就绪决策以及工程品质改进,展现技术领导力和指导能力。
- 开发和评审Kubernetes网络、网络插件、分布式控制平面、Linux主机网络、虚拟机网络和网络自动化方面的生产软件。
- 将产品和基础设施需求转化为安全的编排服务、API、组件边界、状态模型、兼容性策略和交付计划,同时推动共同决策。
- 推动Open vSwitch和OVN在流行为、配置、生命周期管理、升级、互操作性、性能、故障恢复和大规模生产运维中的集成。
- 构建并实施与持续集成、部署和发布验证流程相连的自动化单元测试、集成测试、系统测试、性能测试、扩展性测试和升级测试。
- 参与团队的值班轮班,并与合作伙伴团队共同处理网络升级问题。使用数据包捕获、SDN状态、遥测、分析、受控实验和源码调试来解决事件并防止再次发生。
- 定义并改进可靠性、性能、安全性和资源效率目标,同时为生产网络工程可操作的监控、遥测、追踪和服务级可见性。
我们需要看到:
- 计算机科学或相关领域的学士或硕士学历
查看英文原文
Within theNVIDIA Gaming Network (NGN) infrastructure/storage group, theNGN Software-Defined Storage (SDS) team focuses on storage and infrastructure software for NGN/GeForce NOW, with work spanning NGN 2.x Kubernetes, block storage, performance/scale, security, tenant storage, tenant gameseat, and related central NGN services.
NVIDIA is looking for a Senior System Software Engineer (SDN) to advance software-defined networking for global GPU cloud infrastructure that enables AI workloads, cloud gaming, content delivery, and other accelerated services. Build and operate secure, programmable cloud networking using Open vSwitch, OVN, Kubernetes, C, and Go. Collaborate with senior engineers, NVIDIA architects, and partner teams to lead architecture, implementation, qualification, issue response, and production improvements.
What you'll be doing:
- Collaborate with engineers and architects to define, review, and evolve multi-tenant control-plane and data-plane architecture using Open vSwitch, OVN, OpenFlow, and overlay network technologies.
- Demonstrate technical leadership and mentorship through design and code reviews, implementation guidance, production-readiness decisions, and improvements to engineering quality.
- Develop and review production software for Kubernetes networking, network plugins, distributed control planes, Linux host networking, virtual-machine networking, and network automation.
- Translate product and infrastructure requirements into secure orchestration services, APIs, component boundaries, state models, compatibility strategies, and delivery plans while driving shared decisions.
- Drive Open vSwitch and OVN integration across flow behavior, configuration, lifecycle management, upgrades, interoperability, performance, failure recovery, and large-scale production operation.
- Build and implement automated unit, integration, system, performance, scale, and upgrade tests connected to continuous integration, deployment, and release-qualification workflows.
- Share the team’s on-call rotation and lead networking escalations with partner teams. Use packet captures, SDN state, telemetry, profiling, controlled experiments, and source debugging to resolve incidents and prevent recurrence.
- Define and improve reliability, performance, security, and resource-efficiency objectives while engineering actionable monitoring, telemetry, tracing, and service-level visibility for production networks.
What we need to see:
- BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 12+ years of proven experience designing, implementing, testing, and maintaining production software in both C and Go, with Bash and Python used for test, diagnostic, build, or operational automation.
- Concrete experience developing, integrating, and solving issues in Open vSwitch and OVN, including OpenFlow and control-plane-to-data-plane behavior.
- Production Kubernetes networking experience, including container network interfaces, network policy, node and pod traffic paths, network-plugins, upgrades, and failure modes.
- Strong Linux networking fundamentals and practical understanding of IP, TCP, UDP, routing, switching, overlay networks, tunneling, network namespaces, and network policy.
- Experience architecting distributed systems with secure service APIs, state management, consistency, scalability, compatibility, and failure handling.
- Experience handling the production lifecycle through automated testing, CI/CD, deployments, upgrades, observability, and performance analysis.
- Experience supporting production services through a shared on-call rotation and coordinating incident response through evidence from flows, logs, metrics, traces, profiles, experiments, and source-level debugging.
- Collaborative architecture and technical leadership through written designs, consequential decisions, cross-team work, critical reviews, mentoring, and delivery of significant systems through production operation.
Ways to stand out from the crowd:
- Contributions to Open vSwitch, OVN, OVN-Kubernetes, Kubernetes networking, or another relevant open-source project.
- Experience with large-scale cloud and accelerated virtualization systems, including SR-IOV, RDMA, SmartNICs, data processing units, network function virtualization, KVM/QEMU, or container runtimes.
- Experience building secure, high-performance gRPC or REST services using transport security and strong authentication.
- Applied use of agentic AI and AI-assisted software-developmenttooling, including coding agents, reusable skills, or Model Context Protocol integrations.
With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology industry's most desirable employers. We work with some of the most forward-thinking, versatile people in the world, and our engineering teams are growing fast in some of the most impactful fields of our generation: Cloud Gaming, Cloud Streaming, Network Site Reliability/. If you're a creative engineer who enjoys autonomy and shares our passion for technology, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 17, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.Originally posted on Himalayas