人工智能基础设施与平台运维工程师 (欧盟远程)
AI Infrastructure & Platform Operations Engineer (remote in the EU)
我们正在组建一支欧洲人工智能基础设施与平台运维团队,负责运营由NVIDIA GPU、高性能网络、Kubernetes和下一代平台技术驱动的大规模人工智能基础设施环境。
该团队负责确保在多个数据中心部署的关键人工智能基础设施平台的可用性、性能和运营稳定性。在基础设施、网络和平台运维的交叉领域工作,你将帮助支持推动现代人工智能工作负载的环境。
这是一个有机会接触人工智能基础设施最新技术,并通过k0rdent AI等平台为人工智能驱动的运维服务发展做出贡献的机会。
职责:
- 监控、操作和支持生产环境的人工智能基础设施平台。
- 调查并解决基础设施、网络、硬件和平台相关事件。
- 支持NVIDIA GPU基础设施及相关的平台服务。
- 监控和排查基于Kubernetes的环境。
- 调查基础设施和平台组件中的性能、可用性和可靠性问题。
- 与工程团队、硬件供应商、数据中心人员和服务交付团队协作以解决技术问题。
- 参与事件响应、根本原因分析和运维改进活动。
- 为监控、可观测性、自动化和运维流程的改进做出贡献。
- 维护运维文档、操作手册和知识文章。
- 3年以上基础设施运维、平台运维、网络运维、站点可靠性工程、云运维、数据中心运维或相关技术岗位经验。
- 熟练掌握Linux系统管理与故障排查技能。
- 对网络概念有良好的理解,并具备诊断基础设施相关问题的经验。
- 具备在生产环境中使用Kubernetes的知识。
- 具备支持生产环境基础设施和服务的经验。
- 强大的分析和解决问题的能力。
- 具备在结构化运维和事件管理流程中工作的经验。
- 优秀的沟通和协作能力。
- 能够在轮班制的运维环境中工作。
在以下任一领域有经验者优先考虑:
- NVIDIA GPU基础设施和加速计算平台。
- InfiniBand网络技术
查看英文原文
We are building a European AI Infrastructure & Platform Operations team responsible for operating large-scale AI infrastructure environments powered by NVIDIA GPUs, high-performance networking, Kubernetes, and next-generation platform technologies.
The team is responsible for ensuring the availability, performance, and operational stability of critical AI infrastructure platforms deployed across multiple datacenters. Working at the intersection of infrastructure, networking, and platform operations, you will help support the environments that power modern AI workloads.
This is an opportunity to work with some of the latest technologies in AI infrastructure while contributing to the evolution of AI-powered operational services through platforms such as k0rdent AI.
Responsibilities:
- Monitor, operate, and support production AI infrastructure platforms.
- Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
- Support NVIDIA GPU infrastructure and associated platform services.
- Monitor and troubleshoot Kubernetes-based environments.
- Investigate performance, availability, and reliability issues across infrastructure and platform components.
- Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues.
- Participate in incident response, root cause analysis, and operational improvement activities.
- Contribute to improvements in monitoring, observability, automation, and operational processes.
- Maintain operational documentation, runbooks, and knowledge articles.
- 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
- Strong Linux administration and troubleshooting skills.
- Good understanding of networking concepts and experience diagnosing infrastructure-related issues.
- Working knowledge of Kubernetes in production environments.
- Experience supporting production infrastructure and services.
- Strong analytical and problem-solving skills.
- Experience working within structured operational and incident management processes.
- Excellent communication and collaboration skills.
- Ability to work within a shift-based operational environment.
Experience in one or more of the following areas is highly desirable:
- NVIDIA GPU infrastructure and accelerated computing platforms.
- InfiniBand networking and NVIDIA UFM.
- Kubernetes platform operations.
- AI infrastructure or HPC environments.
- Site Reliability Engineering (SRE) or Platform Engineering.
- Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
- Infrastructure automation technologies and Infrastructure-as-Code practices.
- Large-scale distributed systems and production platforms.
What does Mirantis offer you?
- Work with some of the most advanced AI infrastructure environments in production today.
- Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
- Help define how next-generation AI infrastructure is operated and supported.
- Be part of a team shaping the future of AI-powered operations through k0rdent AI.
- Join a growing organisation investing heavily in AI infrastructure and platform services.
- Salary range: $60000-$67000 gross per year
It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to
By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.
#remote
We are a Leader for Container Management in G2 (#2 after AWS)!
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.
Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen. Learn more at .
Originally posted on Himalayas