高级机器学习工程师(令牌工厂)
Senior ML Engineer (Token Factory)
Nebius简介:
Nebius正在引领全球AI经济的云基础设施新纪元。我们构建了一个全栈AI云平台,支持开发者和企业从数据和模型训练到生产部署的全流程,无需承担构建大型内部AI/ML基础设施的成本和复杂性。
由工程师打造,为工程师服务。从大规模GPU编排到推理优化,我们在计算、存储、网络和应用AI领域掌握各种难题。
在纳斯达克上市(股票代码:NBIS),总部位于阿姆斯特丹,我们在欧洲、英国、北美和以色列设有研发中心,拥有全球化的业务布局。我们的团队超过1500人,其中包括数百名在硬件、软件和AI研发方面具有深厚专业知识的工程师。
职位描述
Token Factory是Nebius Cloud的一部分,这是全球最大的GPU云之一,运行着数以万计的GPU。我们正在构建一个推理与微调平台,使各种基础模型——文本、视觉、音频以及新兴的多模态架构——在大规模训练和部署时变得快速、可靠且易于使用。
我们目前正在进行的一些方向,你可以参与其中:
- 高级微调:提升基于LoRA和全参数的微调方法,用于前沿大语言模型(如GPT-OSS、Kimi K2.5、DeepSeek V3.1/V3.2、GLM-4.7),重点关注模型质量和训练效率。
- 推理优化:识别大语言模型推理中的瓶颈,以提高生产效率。这包括在JAX中构建模型训练和评估流水线,进行推测解码实验,尝试不同的架构(密集型/MoE,自回归/并行),并推导出扩展定律以指导资源分配。
- 低精度训练与推理:研究适用于监督微调和强化学习的低精度(FP8、NVFP4/MXFP4)方法,涵盖推理和训练,并针对现代硬件进行优化。
我们期望你具备:
- 对机器学习和强化学习的理论基础有深刻理解。
- 在语言处理和生成的现代深度学习方面有深入的专业知识。
- 有在多个计算节点上训练大型模型的经验。
- 对大型神经网络训练的性能方面有基本理解(如分片策略、自定义内核、硬件特性等)。
- 强大的软件工程能力
查看英文原文
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
Token Factory is a part of Nebius Cloud, one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference & fine-tuning platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to train & deploy at massive scale.
Some directions we currently working on and which you can be a part of:
·
Advanced Fine-Tuning: Enhancing fine-tuning methodologies - both LoRA-based and full-parameter - for cutting-edge LLMs (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-4.7), focusing on both model quality and training efficiency.
- Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. This involves building model training and evaluation pipelines in JAX for speculative decoding, experimenting with architectures (dense/MoE, auto-regressive/parallel), and deriving scaling laws to guide resource allocation.
- Low Precision Training & Inference: Investigating low-precision (FP8, NVFP4/MXFP4) methodologies for supervised fine-tuning and reinforcement learning - spanning both inference and training - optimized for modern hardware
We expect you to have:
- A profound understanding of theoretical foundations of machine learning and reinforcement learning.
- Deep expertise in modern deep learning for language processing and generation
- Experience with training large models on multiple computational nodes
- Reasonable understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.)
- Strong software engineering skills (we mostly use Python)
- Deep experience with modern deep learning frameworks (we use JAX)
- Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing
- Strong communication and leadership abilities
Nice to have:
- Previous experience working with language models or other similar NLP technologies.
- Familiarity with important ideas in LLM space, such as MHA, RoPE, ZeRO/FSDP, Flash Attention, quantization
- A track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment.
- Strong engineering skills, including experience in developing large distributed systems or high-load web services.
- Open-source projects that showcase your engineering prowess
- Excellent command of the English language, alongside superior writing, articulation, and communication skills.
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
If you need accommodations during the application process, please let us know.