基础设施软件工程师
Infrastructure Software Engineer
关于Nebius:
Nebius正在引领全球AI经济的云基础设施新纪元。我们打造了一个全栈AI云平台,支持开发者和企业从数据和模型训练到生产部署的全流程,无需承担构建大型内部AI/ML基础设施的成本和复杂性。
由工程师打造,为工程师而生。从大规模GPU编排到推理优化,我们在计算、存储、网络和应用AI领域都负责解决难题。
在纳斯达克上市(股票代码:NBIS),总部位于阿姆斯特丹,我们在欧洲、英国、北美和以色列设有研发中心,拥有全球化的业务布局。我们的团队超过1500人,包括数百名在硬件、软件和AI研发方面有深厚经验的工程师。
职位描述
Nebius运营大规模、关键任务的裸金属基础设施。作为软件工程师(Python),你将设计和构建系统,用于大规模地配置、测试和管理物理硬件。你的工作将贴近硬件——直接与服务器、网络和管理控制器交互——同时支持高度自动化、可靠的基础设施运维。
你将与硬件、网络和数据中心运维团队紧密合作,确保我们的平台具备稳健性、可扩展性和生产就绪状态。
你的职责包括:
- 使用Python设计和开发后端服务和自动化系统
- 构建和维护用于硬件配置、测试和生命周期管理的系统
- 开发直接运行在裸金属环境中的软件
- 与Linux系统集成,必要时使用Bash和底层工具
- 实现并维护面向基础设施的CI/CD流水线
- 与网络服务协作,包括IPv4/IPv6、DHCP、DNS、网络启动和服务器启动流程
- 与BMC控制器和管理协议(IPMI风格协议、基于HTTP的标准)进行交互
- 在大规模设备集群中实现可靠的硬件交互和自动化
- 支持ARM64 / ARM64EC架构
- 设计和集成NoSQL数据库用于系统状态和编排数据
- 编写清晰的文档并推动运维卓越
我们期望你具备:
- 作为软件工程师的丰富专业经验,专注于Python
- 熟悉Linux系统和shell脚本
- 有操作裸金属服务器或底层基础设施的实际经验
- 对网络功能有深入理解
查看英文原文
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
Nebius operates large-scale, mission-critical bare-metal infrastructure. As a Software Engineer (Python), you will design and build systems that provision, configure, test, and manage physical hardware at scale. Your work will sit close to the metal—interfacing directly with servers, networks, and management controllers—while supporting highly automated, reliable infrastructure operations.
You will collaborate closely with hardware, networking, and data center operations teams to ensure our platforms are robust, scalable, and production ready.
Your responsibilities will include:
- Design and develop backend services and automation in Python
- Build and maintain systems for hardware provisioning, testing, and lifecycle management
- Develop software that runs directly on bare-metal environments
- Integrate with Linux systems, using Bash and low-level tooling where needed
- Implement and maintain CI/CD pipelines for infrastructure-focused software
- Work with networking services including IPv4/IPv6, DHCP, DNS, network boot, and server boot workflows
- Interface with BMC controllers and management protocols (IPMI-style protocols, HTTP-based standards)
- Enable reliable hardware interaction and automation across large fleets
- Support ARM64 / ARM64EC architectures
- Design and integrate NoSQL data stores for system state and orchestration data
- Write clear documentation and contribute to operational excellence
What we expect you to have:
- Strong professional experience as a software engineer, with a focus on Python
- Solid experience with Linux systems and shell scripting
- Hands-on experience working with bare-metal servers or low-level infrastructure Strong understanding of networking fundamentals (IPv4/IPv6, DHCP, DNS, PXE / network boot)
- Experience interacting with hardware management interfaces (BMC, IPMI-like protocols, HTTP APIs)
- Familiarity with CI/CD systems and production deployment workflows
- Experience designing or working with NoSQL databases
- Ability to debug complex issues spanning software, hardware, and networks
- Strong ownership mindset and clear communication skills in a distributed team
It will be an added bonus if you have:
- Experience operating or building systems for large-scale infrastructure
- Familiarity with ARM-based platforms in production environments
- Background in hardware testing, validation, or factory provisioning
- Experience with infrastructure automation or internal platform tooling
- Contributions to open-source or internal systems software projects
Working conditions:
- Fully remote position (United States)
- Collaboration with globally distributed engineering and operations teams
Key employee benefits:
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families
- 401(k) plan: up to 4% company match with immediate vesting
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers
- Remote work reimbursement: up to $85/month for mobile and internet
- Disability & life insurance: company-paid short-term, long-term, and life insurance coverage
Compensation
- We offer competitive salaries, ranging from $150k- $210k base equity + quarterly performance bonuses.
Join Nebius today and help build the software that powers the next generation of
AI infrastructure.
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
If you need accommodations during the application process, please let us know.