高级支持工程师 - 欧洲
Senior Support Engineer - EU
Nebius正在引领全球人工智能经济的云计算基础设施新时代。我们正在构建一个全栈式的人工智能云平台,支持开发者和企业从数据和模型训练到生产部署的全流程,无需承担自建大型内部AI/ML基础设施的成本和复杂性。
由工程师打造,面向工程师。从大规模GPU编排到推理优化,我们在计算、存储、网络和应用人工智能领域都负责解决难题。
在纳斯达克上市(NBIS),总部位于阿姆斯特丹,我们拥有覆盖欧洲、英国、北美和以色列的全球研发中心。我们的团队超过1500人,其中包括数百名在硬件、软件和人工智能研发方面有深厚专业知识的工程师。
职位描述
我们需要一位高级技术支持工程师,能够处理现代云计算环境中的复杂技术问题。这不是传统意义上的支持岗位。工作内容具有高度的技术性和实操性:调试Linux和Kubernetes问题,调查云计算基础设施中的问题,并帮助客户运行人工智能工作负载、分布式系统和基于GPU的环境。你将与工程团队紧密合作处理生产环境的问题,协助改进内部工具和故障排查流程,并在问题不明确或影响重大时担任升级处理点。
该职位需要轮班值守周末并参与紧急事件响应。
职责包括:
- 调查并解决客户环境中的复杂技术问题
- 在Linux、Kubernetes、云计算基础设施、网络、存储和与GPU相关的任务中进行故障排查
- 支持运行容器化系统、推理任务、训练作业或其他分布式平台的客户
- 作为生产事故的高级升级处理点
- 复现问题,缩小根本原因,并与工程团队合作解决长期问题
- 编写或改进内部脚本、故障排查工具和操作文档
- 通过更好的自动化、可观测性和流程改进使支持工作更具可扩展性
- 在调查和事故期间与客户进行清晰沟通
- 参与周末值班和紧急问题响应
我们寻找的候选人:
- 强大的Linux故障排查技能
- 强大的Kubernetes和容器经验
- 对AWS、GCP、Azure、OpenStack或类似环境中的云计算基础设施有扎实的理解
- 良好的网络基础知识
- 能够用Python、Bash等编写脚本或小型工具
查看英文原文
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
We're looking for a senior support engineer who can handle difficult technical issues in modern cloud environments. This is not a traditional support role. The work is hands-on and technical: debugging Linux and Kubernetes issues, investigating problems in cloud infrastructure, and helping customers running AI workloads, distributed systems, and GPU-based environments. You'll work closely with engineering on production issues, help improve internal tools and troubleshooting workflows, and act as an escalation point when problems are unclear or high impact.
The role includes weekend rotation and incident response.
What you'll do
- Investigate and resolve complex technical issues in customer environments
- Troubleshoot across Linux, Kubernetes, cloud infrastructure, networking, storage, and GPU-related workloads
- Support customers running containerized systems, inference workloads, training jobs, or other distributed platforms
- Act as a senior escalation point for production incidents
- Reproduce issues, narrow down root causes, and work with engineering on long-term fixes
- Build or improve internal scripts, troubleshooting tools, and operational documentation
- Help make support more scalable through better automation, observability, and process improvements
- Communicate clearly with customers during active investigations and incidents
- Take part in weekend coverage and urgent issue response
What we're looking for
- Strong Linux troubleshooting skills
- Strong Kubernetes and container experience
- Solid understanding of cloud infrastructure in AWS, GCP, Azure, OpenStack, or similar environments
- Good networking fundamentals
- Ability to write scripts or small tools in Python, Bash, Go, or similar
- Experience working on production issues that require structured debugging and cross-team collaboration
- Ability to work independently and stay effective when the path to resolution is not obvious
- Clear written communication, especially when explaining technical issues to customers and internal teams
Especially valuable
- Experience with GPU-based infrastructure
- Familiarity with AI/ML or LLM-related workloads
- Understanding of inference and training pipelines
- Experience improving observability, tooling, or operational workflows
- History of building useful internal tools or automating repetitive work
- Personal or open-source projects that show real technical depth
A strong candidate for this role usually
- enjoys debugging messy infrastructure problems
- takes ownership without waiting to be told exactly what to do
- thinks beyond the immediate ticket
- works well with engineering
- looks for ways to reduce repeated operational pain, not just close cases
#LI-VD1
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
If you need accommodations during the application process, please let us know.