远程工作雷达

高级解决方案工程师 – GPU 与 AI 基础设施

Senior Solution Engineer – GPU & AI Infrastructure

AI开发工程市场运营限定地区(需当地身份)日间重叠仅 1 小时,需熬夜配合
公司Civo
薪资未公开
工作地点United Kingdom
地域资格限定地区(需当地身份)
时区要求日间重叠仅 1 小时,需熬夜配合
用工类型permanent
发布时间2026-08-11
数据来源4dayweek.io
前往 4dayweek.io 查看并投递 →
注意地域限制:该职位明确限定在 United Kingdom 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。
作息提示:日间重叠仅 1 小时,需熬夜配合。

#### **高级解决方案工程师 – GPU 与 AI 基础设施**

* * *

**关于 Civo:**

Civo 是一家高性能的新云服务提供商,专为现代 AI、高性能计算(HPC)和云原生基础设施的需求而设计。我们消除传统云的冗余开销,提供超低延迟计算、裸金属 GPU 性能以及可扩展的 Kubernetes 编排。

专为 AI 工程团队、企业及研究机构设计,Civo 提供直接访问前沿 NVIDIA GPU 集群、高速网络架构和并行存储系统的权限,以高效地训练、微调和部署基础模型。我们将高密度基础设施与可预测的定价和最大计算吞吐量相结合,使组织能够在不增加复杂性或成本负担的情况下扩展 AI 工作负载。

* * *

**职位描述:**

作为高级解决方案工程师 – GPU 与 AI 基础设施,您将担任 Civo 大规模 AI 和高性能计算(HPC)客户项目的首要技术架构师。您将负责设计用于训练和推理大规模基础模型的先进 NVIDIA GPU 集群。

在该职位中,您将弥合客户业务目标与超高性能硬件执行之间的差距。您将主导技术对接,将复杂的 AI 工作负载需求转化为生产就绪的高层设计(HLD)、低层设计(LLD)和详细物料清单(BOM)。您的专业知识将涵盖基于裸金属和 Kubernetes 的编排,适用于前沿 NVIDIA Blackwell 架构(如 B300 和 GB300NVL),使用超低延迟 InfiniBand 和高速 RoCE 网络架构。

* * *

**职责:**

**解决方案设计与架构**

- 系统设计文档:撰写企业级 GPU 超算集群的全面高层设计(HLD)和低层设计(LLD)文档。

- 物料清单(BOM):生成涵盖计算节点、NVLink 交换机、网络架构、收发器/线缆、液冷/风冷需求、电力分配和高性能存储的详细 BOM。

- GPU 集群拓扑:为 NVIDIA Blackwell 平台(特别是 B300 和 GB300NVL 机架级架构)设计扩展型(NVLink/NVSwitch)和扩展外(Fat-Tree、Rail-Optimized)网络拓扑。

- 网络架构工程:设计高吞吐、低延迟的网络架构。

查看英文原文

#### **Senior Solution Engineer – GPU & AI Infrastructure**

* * *

**About Civo:**

Civo is a high-performance neocloud provider purpose-built for the demands of modern AI, high-performance computing (HPC), and cloud-native infrastructure. We eliminate legacy cloud overhead to deliver ultra-low-latency compute, bare-metal GPU performance, and streamlined Kubernetes orchestration at scale.

Purpose-designed for AI engineering teams, enterprises, and research institutions, Civo delivers direct access to cutting-edge NVIDIA GPU clusters, high-speed fabrics, and parallel storage systems required to train, fine-tune, and deploy foundation models efficiently. We combine high-density infrastructure with predictable pricing and maximum compute throughput, empowering organizations to scale AI workloads without the complexity or cost bloat of traditional hyperscalers.

* * *

**About the Role:**

As a Senior Solution Engineer – GPU & AI Infrastructure, you will serve as the primary technical architect for Civo’s large-scale AI and high-performance computing (HPC) customer initiatives. You will be responsible for designing state-of-the-art NVIDIA GPU clusters tailored for training and inferencing massive foundation models.

In this role, you will bridge the gap between customer business objectives and ultra-high-performance hardware execution. You will lead technical engagements, translate complex AI workload requirements into production-ready High-Level Designs (HLD), Low-Level Designs (LLD), and detailed Bills of Materials (BOM). Your expertise will span bare-metal and Kubernetes-based orchestrations across cutting-edge NVIDIA Blackwell architectures (e.g., B300 and GB300NVL) using ultra-low-latency InfiniBand and high-speed RoCE networking fabrics.

* * *

**Responsibilities:**

**Solution Design & Architecture**

- System Design Documents: Author comprehensive High-Level Design (HLD) and Low-Level Design (LLD) documentation for enterprise-scale GPU supercomputing clusters.

- Bill of Materials (BOM): Generate detailed BOMs covering compute nodes, NVLink switches, network fabrics, transceivers/cabling, liquid/air cooling requirements, power distribution, and high-performance storage.

- GPU Cluster Topology: Architect scale-up (NVLink/NVSwitch) and scale-out network topologies (Fat-Tree, Rail-Optimized) for NVIDIA Blackwell platforms, specifically B300 and GB300NVL rack-scale architectures.

- Fabric & Networking Engineering: Design high-throughput, low-latency networking architectures utilizing both InfiniBand (e.g., NDR/X800) and RoCE / RoCEv2 (e.g., NVIDIA Spectrum-X / Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing).

- Multi-Tenant & Deployment Models: Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native / Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, KubeFlow).

- Storage Integration: Architect high-bandwidth parallel storage solutions utilizing GPUDirect Storage (GDS) and enterprise AI file systems (e.g., VAST Data).

**Technical Sales Support & Customer Engagement**

- Partner with Civo’s sales and commercial teams as the technical lead for high-value AI infrastructure opportunities.

- Engage directly with customer CTOs, Chief AI Officers, infrastructure leads, and ML engineers to evaluate technical requirements, compute sizing, and fabric choices.

- Lead deep-dive architectural workshops and technical presentations on Civo's bare-metal GPU and managed Kubernetes offerings.

- Produce precise technical proposals and lead responses to complex RFPs/RFIs regarding AI infrastructure.

**Proof-of-Concept (PoC) & Benchmarking**

- Architect and oversee Proof-of-Concept (PoC) deployments to validate real-world performance for customer workloads.

- Benchmark cluster performance using industry-standard tools (NCCL tests, GPUDirect RDMA latency/bandwidth, MLPerf, Megatron-LM benchmarks).

- Address network congestion, fabric routing, and thermal/power optimization during validation phases.

**Product & Ecosystem Collaboration**

- Serve as the bridge between enterprise AI clients, hardware vendors (NVIDIA, network OEMs), and Civo’s internal platform engineering team.

- Provide continuous feedback to product teams on market trends, hardware platform demands, and feature requirements for AI/GPU orchestration.

* * *

**Key Results/Objectives:**

- Technical Wins: Achieve high technical win rates on large-scale AI/GPU cluster sales opportunities.

- Design Excellence: Successfully deliver complete, peer-reviewed HLDs, LLDs, and BOMs within target deal timelines.

- Customer Satisfaction: Achieve successful PoC completion and sign-off for enterprise clients scaling AI workloads on Civo infrastructure.

* * *

**Requirements:**

#### **Experience & Core Qualifications**

- 5+ years in a Solution Architecture, Systems Engineering, or Technical Pre-Sales role focused on high-performance cloud, HPC, or AI infrastructure.

- Bachelor’s degree in Computer Science, Electrical Engineering, Systems Engineering, or equivalent practical experience.

#### **Technical Expertise**

- NVIDIA GPU Architecture: Deep hands-on knowledge of NVIDIA HGX/DGX platforms, NVLink/NVSwitch fabrics, and Blackwell architectures (B300, GB300NVL, GB200 NVL72/NVL36).

- High-Speed Networking: Expert-level knowledge of cluster fabric topologies:

- InfiniBand: Quantum-2 / Quantum-X800, Subnet Management, Adaptive Routing.

- RoCE / RoCEv2: Spectrum-X / Spectrum-4 Ethernet switches, PFC, ECN, RoCE configuration, and optimization.

- GPU Direct Technologies: GPUDirect RDMA (GDR) and GPUDirect Storage (GDS).
- Orchestration & Platforms: Proficiency in deploying and optimizing GPU workloads on:

- Kubernetes: Container networking (CNI), NVIDIA GPU Operator, RDMA Shared Device Plugin, MPI Operator.

- Bare-Metal: Slurm, Ansible, Terraform, PyTorch/NCCL environment tuning.
- Documentation Skills: Demonstrated experience creating enterprise-grade HLDs, LLDs, network rack diagrams, and itemized BOMs.

- Power & Thermal Awareness: Familiarity with high-density datacenter environments, liquid cooling technologies (Direct-to-Chip, CDU/liquid loop setups), and power delivery constraints for 100kW+ per rack deployments.

#### **Soft Skills**

- Strong technical leadership and presentation skills, with the ability to articulate complex network and hardware tradeoffs to executive stakeholders.

- Problem-solving mindset capable of diagnosing complex hardware-software interaction bottlenecks in distributed training/inference setups.

#### **Location**

- Must be UK based.

* * *

**Nice to Have:**

- NVIDIA Certified Professional: AI Infrastructure (NCP-AII).

- NVIDIA Certified Professional: AI Networking (NCP-AIN).

- NVIDIA Certified Professional: InfiniBand (NCP-IB).

- NVIDIA Certified Associate / Professional: AI Workload Deployment & Cloud Native.

* * *

**Why Join Civo?**

- Competitive compensation and benefits package.

- 4-day week company (unless attending an event).

- Uncapped holiday.

- Remote work environment with flexibility and autonomy.

- Collaborative and inclusive culture that values diversity and creativity.

- Opportunity to work with a dynamic and innovative team in the fast-growing cloud industry.

本页面信息整理自 4dayweek.io,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

销售开发代表

CivoIndiapermanent2026-04-29
市场运营限定地区(需当地身份)

← 返回全部职位