高级ML解决方案架构师 - 代币工厂
Senior ML Solutions Architect - Token Factory
关于Nebius:
Nebius正在引领全球AI经济的云基础设施新纪元。我们正在构建一个全栈AI云平台,支持开发者和企业从数据和模型训练到生产部署,无需承担构建大型内部AI/ML基础设施的成本和复杂性。
由工程师打造,面向工程师。从大规模GPU编排到推理优化,我们在计算、存储、网络和应用AI方面都解决关键难题。
在纳斯达克上市(NBIS),总部位于阿姆斯特丹,我们在欧洲、英国、北美和以色列设有研发中心,拥有全球业务布局。我们的团队超过1500人,包括数百名在硬件、软件和AI研发方面有深厚专业知识的工程师。
职位描述
该职位位于Nebius Token Factory,这是我们用于在生产环境中运行和定制开源大语言模型的无服务器平台。Token Factory提供无服务器推理和微调(LoRA、完整微调、RFT),并由内部优化技术如自定义推测解码、量化、缓存感知路由和专用端点支持。客户选择我们是为了从原型过渡到可扩展的生产环境,而无需承担构建和调整自己推理堆栈的成本和复杂性。
我们寻求一位经验丰富的高级机器学习解决方案架构师,以支持客户利用Nebius Token Factory的无服务器推理和微调平台,在多种模态上使用开源大语言模型。在此职位中,您将与客户合作,设计和实现优化的推理工作流,构建定制的基于大语言模型的解决方案,并使用我们提供的模型构建可扩展的AI应用。您还将与我们的后端团队紧密合作,改进我们的平台以满足客户的需求。
您可以在欧洲远程办公。
您的职责将包括:
- 在各种模态上优化大语言模型推理,以推动业务价值并支持客户目标
- 提供监督学习和强化学习微调支持,以最大化客户的模型质量
- 使用Nebius Token Factory的推理服务设计和实现基于大语言模型的解决方案
- 构建利用我们无服务器大语言模型API的生产就绪应用程序,包括多模态模型(文本、视觉、音频)和领域特定模型
- 提供提示工程、RAG架构和模型选择方面的技术专长
- 与产品和工程团队合作,收集客户反馈并塑造平台
查看英文原文
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
This position sits within Nebius Token Factory, our serverless platform for running and customizing open-source LLMs in production. Token Factory allows for serverless inference and fine-tuning (LoRA, full FT, RFT) backed by in-house optimizations like custom speculative decoding, quantization, cache-aware routing and dedicated endpoints. Customers come to us to move from prototype to scaled production without the cost and complexity of building and tuning their own inference stack.
We seek an experienced Senior ML Solutions Architect to support customers leveraging Nebius Token Factory's serverless inference and fine-tuning platforms for open-source LLMs across multiple modalities. In this role, you will be collaborating with clients to design and implement optimized inference workflows, build customized LLM-based solutions and architect scalable AI applications using our served models. You will also work closely with our backend team to improve our platform to match clients' needs.
You’re welcome to work remotely from Europe.
Your responsibilities will include:
- Optimize LLM inference across various modalities to drive business value and support customer goals
- Provide support in supervised and reinforcement learning fine-tuning to maximize model quality for the customers
- Design and implement LLM-based solutions using Nebius Token Factory’s inference services
- Build production-ready applications leveraging our serverless LLM APIs, including multimodal models (text, vision, audio) and domain-specific models
- Provide technical expertise in prompt engineering, RAG architectures and model selection
- Collaborate with product and engineering teams to surface customer feedback and shape the platform roadmap
- Guide customers in scaling from POC to production with a focus on performance, reliability, and cost efficiency
We expect you to have:
- 5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI
- Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches
- Hands-on experience with:
- Running LLMs in production: deploying and operating inference workloads
- LLM fine-tuning, including supervised fine-tuning (SFT/LoRA) and data preparation/curation; experience with RL-based fine-tuning is a strong plus
- LLM evaluation: building task-specific benchmarks and offline/online eval pipelines, including LLM-as-a-judge setups
- Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers)
- Deploying LLM-powered applications using APIs from OpenAI, Anthropic, or open-source models
- Strong Python programming skills
- Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences
It would be an added bonus if you have:
- Work with multimodal AI models (e.g., vision-language, speech)
- Proficiency with DevOps tools (Docker, Kubernetes)
- Contributions to open-source ML/AI projects
Preferred technical stack:
- Programming Languages: Python
- ML Frameworks and Libraries: vLLM, TensorRT-LLM, SGLang, Transformers, OpenAI/Anthropic SDKs
- MLOps and DevOps tools: Kubernetes (K8s), Docker, Git
- Cloud Platforms: AWS (SageMaker, Bedrock), GCP (Vertex AI), Azure (Azure ML)
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
If you need accommodations during the application process, please let us know.
Originally posted on Himalayas