高级机器学习应用与编译器工程师
Senior Machine Learning Applications and Compiler Engineer
我们正在寻找一位高级机器学习应用与编译器工程师!
NVIDIA 正在寻找工程师,为我们的 LPX 推理和编译器堆栈开发算法和优化。你将在大规模系统、编译器和深度学习的交汇点工作,设计神经网络工作负载如何映射到未来的 NVIDIA 平台。这是你参与极具创新性项目的机会!
你将负责:
- 构建、开发和维护高性能运行时和编译器组件,专注于端到端推理优化。
- 定义并实现大规模推理工作负载在 NVIDIA 系统上的映射。
- 扩展并集成到 NVIDIA 的软件生态系统中,为库、工具和接口做出贡献,使模型在不同平台上的部署更加顺畅。
- 对关键性能和效率指标进行基准测试、分析和监控,确保编译器能生成高效的神经网络图到推理硬件的映射。
- 与硬件架构师和设计团队紧密合作,反馈软件观察结果,影响未来架构,并共同设计能提升性能和效率的新功能。
- 原型化并评估新的编译和运行时技术,包括针对空间处理器的图转换、调度策略和内存/布局优化。
- 在顶级机器学习、编译器和计算机架构会议上发表和展示关于推理及相关空间加速器的新型编译方法的技术成果。
我们需要看到:
- 计算机科学、电子/计算机工程或相关领域的硕士或博士学位,或同等经验,具有 5 年相关工作经验。
- 强大的软件工程背景,精通系统级编程(如 C/C++ 和/或 Rust),并具备扎实的数据结构、算法和并发方面的计算机科学基础。
- 具有编译器或运行时开发的实际经验,包括中间表示设计、优化过程或代码生成。
- 具有 LLVM 和/或 MLIR 的经验,包括构建自定义传递、方言或集成。
- 熟悉深度学习框架如 TensorFlow 和 PyTorch,并有使用 ONNX 等可移植图格式的经验。
- 对并行和异构计算架构有深入理解,例如 GPU、空间加速器或其他领域专用处理器。
- 强大的分析和调试能力,有相关经验。
查看英文原文
We are now looking for a Senior Machine Learning Applications and Compiler Engineer!
NVIDIA is seeking engineers to develop algorithms and optimizations for our LPX inference and compiler stack. You will work at the intersection of large-scale systems, compilers, and deep learning, crafting how neural network workloads map onto future NVIDIA platforms. This is your chance to be part of something outstandingly innovative!
What you’ll be doing:
- Build, develop, and maintain high-performance runtime and compiler components, focusing on end-to-end inference optimization.
- Define and implement mappings of large-scale inference workloads onto NVIDIA’s systems.
- Extend and integrate with NVIDIA’s SW ecosystem, contributing to libraries, tooling, and interfaces that enable seamless deployment of models across platforms.
- Benchmark, profile, and monitor key performance and efficiency metrics to ensure the compiler generates efficient mappings of neural network graphs to our inference hardware.
- Collaborate closely with hardware architects and design teams to feedback software observations, influence future architectures, and codesign features that unlock new performance and efficiency points.
- Prototype and evaluate new compilation and runtime techniques, including graph transformations, scheduling strategies, and memory/layout optimizations tailored to spatial processors.
- Publish and present technical work on novel compilation approaches for inference and related spatial accelerators at top tier ML, compiler, and computer architecture venues.
What we need to see:
- MS or PhD in Computer Science, Electrical/Computer Engineering, or related field, or equivalent experience, with 5 years of relevant experience.
- Strong software engineering background with proficiency in systems level programming (e.g., C/C++ and/or Rust) and solid CS fundamentals in data structures, algorithms, and concurrency.
- Hands on experience with compiler or runtime development, including IR design, optimization passes, or code generation.
- Experience with LLVM and/or MLIR, including building custom passes, dialects, or integrations.
- Familiarity with deep learning frameworks such as TensorFlow and PyTorch, and experience working with portable graph formats such as ONNX.
- Solid understanding of parallel and heterogeneous compute architectures, such as GPUs, spatial accelerators, or other domain specific processors.
- Strong analytical and debugging skills, with experience using profiling, tracing, and benchmarking tools to drive performance improvements.
- Excellent communication and collaboration skills, with the ability to work across hardware, systems, and software teams.
- Ideal candidates will have direct experience with MLIR based compilers or other multilevel IR stacks, especially in the context of graph based deep learning workloads.
Ways to stand out from the crowd:
- Prior work on spatial or dataflow architectures, including static scheduling, pipeline parallelism, or tensor parallelism at scale.
- Contributions to opensource ML frameworks, compilers, or runtime systems, particularly in areas related to performance or scalability.
- Demonstrated research impact, such as publications or presentations at conferences like PLDI, CGO, ASPLOS, ISCA, MICRO, MLSys, NeurIPS, or similar.
- Experience with large-scale AI distributed inference or training systems, including performance modeling and capacity planning for multi rack deployments.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 135,000 CAD - 185,000 CAD for Level 3, and 170,000 CAD - 220,000 CAD for Level 4.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until March 27, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
Originally posted on Himalayas