远程工作雷达

高级AI基础设施工程师(GPU)- 远程欧洲、中东和非洲

AI Infrastructure Engineer (GPU) - Remote EMEA

AI开发工程全球可投(据职位描述推断)
公司pragmatike
薪资未公开
工作地点Ukraine / Czech Republic / Bulgaria / Latvia / Spain / Hungary / Albania / Lithuania / Greece / Bosnia & Herzegovina / Croatia / Estonia / Serbia / Dubai / Poland / Armenia / Portugal / Italy / Malta / Türkiye / Montenegro / Romania
地域资格全球可投(据职位描述推断)
时区要求无特别要求
用工类型FullTime
发布时间2026-08-13
数据来源Ashby
前往企业招聘页投递 →
全球可投:该职位未限制候选人所在地区。仍需注意薪资可能按地区折算,以及实际签约方式(正式雇佣 / 独立合同)。

位置:完全远程(EMEA时区)
开始日期:尽快
语言:需要流利的英语
行业:云计算 / 人工智能 / 欧洲深度科技SaaS

关于该职位

Pragmatike正在为一家快速扩展、资金充足的分布式云基础设施初创公司招聘,该公司正在构建下一代AI原生云服务。这家公司正在重新定义计算的交付方式,通过去中心化架构提供GPU驱动的基础设施用于AI/ML工作负载、安全存储和高速数据传输,与传统云服务商相比显著降低环境影响。

我们正在寻找一位具有丰富经验的AI基础设施工程师,专注于生产级模型服务和AI系统的基础设施。这是一个高度技术性、动手能力强的职位,专注于构建可扩展、可靠且高效的ML推理平台,支持实时AI应用。

你将负责设计和运营服务于大规模机器学习模型的核心基础设施。你将与基础设施、平台和应用AI团队紧密合作,确保高可用性、低延迟和成本高效的推理系统。强烈的所有权意识、生产思维以及对分布式GPU系统的经验是必不可少的。

你的职责

- 使用vLLM、TGI、Triton或类似框架构建和运营生产级模型服务基础设施

- 设计并实现具有蓝/绿和灰度发布策略的稳健部署流水线

- 开发和维护自动扩展系统、多模型服务架构和智能请求路由层

- 优化GPU利用率、内存效率、网络吞吐量和模型构件存储性能

- 设计可观测系统,用于跟踪推理延迟、吞吐量、GPU使用情况、成本指标和系统健康状况

- 管理模型注册表和CI/CD流水线,实现自动化和可重复的模型部署

- 负责ML系统的全生命周期,从开发到生产,包括运维支持和值班责任

- 定义工程最佳实践,并在快速发展的初创环境中为平台可扩展性做出贡献

所需资格

- 4年以上ML Ops、平台工程、SRE或类似专注于ML系统的基础设施相关工作经验

- 具有vLLM、TGI、Triton或类似模型服务框架的实际操作经验

- 强烈的...

查看英文原文

Location: Fully remote (EMEA timezone)
Start date: ASAP
Languages: Fluent English required
Industry: Cloud Computing / AI / European Deep-Tech SaaS

ABOUT THE ROLE

Pragmatike is recruiting on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup building next-generation AI-native cloud services. The company is redefining how compute is delivered by providing GPU-powered infrastructure for AI/ML workloads, secure storage, and high-speed data transfer through a decentralized architecture that significantly reduces environmental impact compared to traditional cloud providers.

We are seeking a AI Infrastructure Engineer with strong experience in production-grade model serving and infrastructure for AI systems. This is a highly technical, hands-on role focused on building scalable, reliable, and efficient ML inference platforms powering real-time AI applications.

You will be responsible for designing and operating the core infrastructure that serves machine learning models at scale. You will work closely with infrastructure, platform, and applied AI teams to ensure high availability, low latency, and cost-efficient inference systems. Strong ownership, production mindset, and experience with distributed GPU systems are essential.

YOUR RESPONSIBILITIES

- Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent

- Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models

- Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers

- Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance

- Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health

- Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments

- Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities

- Define engineering best practices and contribute to platform scalability in a fast-moving startup environment

REQUIRED QUALIFICATIONS

- 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems

- Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent

- Strong background in container orchestration and operating GPU-based workloads in production

- Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines

- Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)

- Strong understanding of distributed systems, performance tuning, and production reliability engineering

- Ability to effectively use AI coding assistants to accelerate development and debugging workflows

- Ownership mindset with the ability to operate independently in a remote-first environment

PREFERRED QUALIFICATIONS

- Experience with ML platforms such as Kubeflow, MLflow, or KubeAI

- Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems

- Experience with cost optimization across different GPU types and inference workloads

- Background in early-stage startups or greenfield infrastructure projects

- Proven experience building production systems from scratch rather than maintaining legacy platforms

WHY JOIN US

- Take ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform

- Build foundational ML inference systems from the ground up in a high-growth, well-funded startup

- Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture

- Gain deep expertise in next-generation AI infrastructure and large-scale model serving systems

- Influence core engineering decisions and define best practices that will scale with the company.

Pragmatike is committed to a fair, transparent, and inclusive recruitment process. We do not discriminate based on age, disability, gender, gender identity or expression, marital or civil partner status, pregnancy or maternity, race, religion or belief, sex, or sexual orientation.

In accordance with GDPR, your personal data will be processed lawfully, fairly, and securely, and used solely for recruitment purposes, including sharing it with our client(s) for employment consideration.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位