远程工作雷达

机器学习基础设施工程师

ML Infrastructure Engineer

开发工程限定地区(需当地身份)
公司nebius
薪资未公开
工作地点Amsterdam, Netherlands; Remote - Europe; Remote - United States
地域资格限定地区(需当地身份)
时区要求无特别要求
用工类型未标注
发布时间2026-05-12
数据来源Greenhouse
前往企业招聘页投递 →
注意地域限制:该职位明确限定在 Amsterdam, Netherlands; Remote - Europe; Remote - United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

Nebius简介:

Nebius正在引领全球AI经济的云基础设施新纪元。我们打造了一个全栈AI云平台,支持开发者和企业从数据和模型训练到生产部署的全流程,无需承担构建大型内部AI/ML基础设施的成本和复杂性。

由工程师打造,面向工程师。从大规模GPU编排到推理优化,我们在计算、存储、网络和应用AI领域掌握着核心难题。

在纳斯达克上市(NBIS),总部位于阿姆斯特丹,我们在欧洲、英国、北美和以色列设有研发中心,拥有全球化的业务布局。我们的团队超过1500人,其中包括数百名在硬件、软件和AI研发方面具有深厚专业知识的工程师。

职位描述

我们正在寻找一位技术精湛的ML/AI工程师加入我们的团队,负责GPU平台在机器学习和AI工作负载方面的基准测试。你将在评估基于GPU的硬件性能方面发挥关键作用,为平台优化和下一代硬件开发提供数据驱动的决策支持。

你的职责包括:

  • 与硬件和开发团队紧密合作,对GPU在系统和内核层面的性能进行分析和剖析。
  • 在不同平台、架构和软件堆栈(如CUDA、ROCm)之间评估和比较GPU性能。
  • 调试和优化ML工作负载,使其在GPU硬件上高效运行,识别并解决性能瓶颈。
  • 对新的GPU集群进行验收测试,确保硬件和软件满足AI工作负载的性能、稳定性和兼容性要求。
  • 在多种GPU系统配置中进行实验,评估不同互连策略和系统级优化对性能和可扩展性的影响。
  • 开发工具和仪表板,用于可视化性能指标、瓶颈和趋势。
  • 参与内部工具、框架和最佳实践的建设

我们期望你具备:

  • 对机器学习理论基础有深刻理解
  • 深入了解大规模神经网络训练和推理的性能方面(数据/张量/上下文/专家并行、卸载、自定义内核、硬件特性、注意力优化、动态批处理等)
  • 在现代深度学习框架上有丰富的经验
查看英文原文

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role

We are seeking a highly skilled ML/AI Engineer to join our team to lead and support benchmarking of GPU platforms benchmarking of GPU platforms for machine learning and AI workloads. You will play a critical role in evaluating the performance of GPU-based hardware for various deep learning and AI frameworks, enabling data-driven decisions for platform optimisation and next-generation hardware development.

Your responsibilities will include:

  • Work closely with hardware, development teams to profile and analyse GPU performance at the system and kernel level.
  • Evaluate and compare GPU performance across different platforms, architectures, and software stacks (e.g.,CUDA, ROCm).
  • Debug and optimise ML workloads to run efficiently on GPU hardware, identifying and resolving performance bottlenecks.
  • Perform acceptance testing acceptance testing for new GPU clusters, ensuring hardware and software meet performance, stability, and compatibility requirements for AI workloads.
  • Perform experiments across diverse GPU system configurations to assess the impact of varying interconnect strategies and system-level optimisations on performance and scalability.
  • Develop tools and dashboards to visualise performance metrics visualise performance metrics, bottlenecks, and trends.
  • Contribute to internal tooling, frameworks, and best practices

We expect you to have:

  • A profound understanding of theoretical foundations of machine learning
  • Deep understanding of performance aspects of large neural networks training and inference (data/tensor/context/expert parallelism, offloading, custom kernels, hardware features, attention optimisations, dynamic batching etc.)
  • Deep experience with modern deep learning frameworks (PyTorch, JAX, Megatron-LM, Tensort-LLM)
  • Good understanding of the GPU stack: CUDA,NCCL, drivers, and relevant libraries
  • Familiarity with containerized environments (e.g., Docker, Kubernetes).
  • Strong communication and ability to work independently

Ways to stand out from the crowd:

  • Familiarity with modern LLM inference frameworks (vLLM, SGLang, TensorRT)
  • Experience in Python and performance profiling tools (e.g., Nsight, nvprof, perf).
  • Familiarity with cloud ML platforms like AWS, GCP, Azure ML
  • Contributions to open-source ML benchmarking tools

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.

If you need accommodations during the application process, please let us know.

本页面信息整理自 Greenhouse,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

定价总监

nebiusRemote2026-06-26
职能支持全球可投

应用安全工程师

nebiusIsrael€75,000 - €240,000/年Full Time今天
开发工程限定地区(需当地身份)

← 返回全部职位