站点可靠性工程师
Site Reliability Engineer
关于Nebius:
Nebius正在引领全球AI经济的云基础设施新纪元。我们正在构建一个全栈AI云平台,支持开发者和企业从数据和模型训练到生产部署的全流程,而无需承担构建大型内部AI/ML基础设施的成本和复杂性。
由工程师打造,面向工程师。从大规模GPU编排到推理优化,我们在计算、存储、网络和应用AI领域都负责解决难题。
在纳斯达克上市(NBIS),总部位于阿姆斯特丹,我们在欧洲、英国、北美和以色列拥有全球研发中心。我们的团队超过1500人,包括数百名在硬件、软件和AI研发方面有深厚专业知识的工程师。
职位描述
Nebius正在寻找一名硬件基础设施团队的站点可靠性工程师。欢迎你在阿姆斯特丹的办公室工作。
硬件基础设施团队负责设计、开发和支持数据中心生命周期中的系统:
- 服务功能和负载测试系统
- 监控位于我们数据中心的工程设备(电源、空气和水冷系统等)
- 监控IT设备:机架、服务器、JBOD、JBOG、电源架、网络设备等
- 资产追踪
- 硬件维修任务追踪
- 服务器生产
在此职位中,你的职责将包括:
- 确保服务的容错性、可扩展性和不间断运行
- 使用前沿技术解决各种基础设施问题
- 实现并改进CI/CD流程
我们期望你具备:
- 精通Linux系统,具备使用Python和Bash脚本进行自动化的经验
- 具备排查复杂系统问题的能力,包括硬件、软件和网络问题
- 强大的分析和解决问题能力,专注于优化系统性能
- 英语工作水平
如果你具备以下条件,将是一个加分项:
- 有兴趣参与后端开发
- 具备设计、开发和运行高负载分布式系统的经验
工作条件:
- 主要远程办公
- 偶尔需要前往数据中心,尤其是如果你不在附近的话
- 与分布在全球的工程和运维团队合作
关键员工福利:
- 健康保险:公司全额支付员工及家庭的医疗、牙科和视力保险
- 401(k)计划:公司最高匹配4%
查看英文原文
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to work in our office in Amsterdam.
Hardware Infrastructure team designs, develops and supports systems involved in the data-centers lifecycle:
- Serving functional and load testing system.
- Monitoring of engineering equipment located in our data centers (power supply, air and water cooling, etc.)
- Monitoring of IT equipment: racks, servers, JBODs, JBOGs, power shelves, network devices, etc.
- Asset tracking.
- Hardware repairs tasks tracking.
- Server production.
In this position, your responsibility will be to:
- Ensure fault-tolerance, scale and uninterrupted operations for our services.
- Use cutting-edge technology to solve a variety of infrastructure problems.
- Implement and improve CI/CD processes.
We expect you to have:
- Proficiency in Linux systems, with expertise in Python and Bash scripting for automation.
- Demonstrated ability to troubleshoot complex system issues, including hardware, software and networking problems.
- Strong analytical and problem-solving skills, with a focus on optimizing system performance.
- Working proficiency in English.
It would be an added bonus if you had:
- Desire to be involved in backend development.
- Experience designing, developing and running high-load distributed systems.
Working conditions:
- Primarily remote
- Occasional travel to data centers required, especially if not located near one
- Collaboration with globally distributed engineering and operations teams
Key employee benefits:
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families
- 401(k) plan: up to 4% company match with immediate vesting
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers
- Remote work reimbursement: up to $85/month for mobile and internet
- Disability & life insurance: company-paid short-term, long-term, and life insurance coverage
Compensation
- We offer competitive salaries, ranging from $130k- $180k base + quarterly performance bonuses.
Join Nebius and help operate the systems that power next-generation AI
infrastructure.
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
If you need accommodations during the application process, please let us know.