AI推理工程师
AI Inference Engineer
Fuse Energy是一家具有前瞻性思维的可再生能源初创公司,致力于快速提供一太瓦的可再生能源。我们结合基本原理思考与前沿技术,打造一个更加优越的能源系统。我们从包括Multicoin、Balderton、Lakestar、Accel、Creandum、Lowercarbon、Ribbit、Box Group在内的顶级投资机构以及Nico Rosberg(Solana联合创始人)等战略天使投资人处筹集了2.1亿美元资金。
随着数据中心成为电力需求最大且增长最快的来源之一,Fuse正在扩展高性能计算基础设施,该领域位于能源和AI的交汇点。我们同时从零开始构建GPU/CUDA性能层和推理服务层,并正在寻找一位创始工程师来负责后者。
我们正在寻找一名创始AI推理工程师,负责定义并构建Fuse如何大规模地提供AI推理工作负载,直接向CTO汇报。与我们的CUDA和GPU工程招聘人员负责内核级和硬件性能不同,该职位负责其上层:模型实际如何被提供、扩展和交付以满足承诺的性能目标。
机会
Fuse在我们运营的市场中对数据中心容量有显著需求,主要集中在推理方面。世界上很少有公司能像Fuse一样将真实的电力输送与真实的计算能力相结合,这使得推理服务成为我们将这一优势转化为市场上最佳产品核心的关键。这就是这个职位。
职责
- 从基本原理出发,定义Fuse的推理服务策略和架构。
- 设计并构建服务堆栈:针对高吞吐量、低延迟的推理工作负载进行请求路由、批处理、调度和自动扩展。
- 负责推理的模型级优化策略——决定在哪里以及如何应用量化、蒸馏、推测解码等技术以提高吞吐量和每token成本,与CUDA/GPU工程师合作。
- 在服务框架和编排方面做出核心软件架构决策(例如vLLM、TensorRT-LLM、SGLang、Triton Inference Server或类似工具)。
- 将吞吐量、延迟和可用性承诺转化为具体的工程技术规范和服务容量计划。
- 直接负责推理性能和可靠性的技术主导。
- 与CUDA和GPU工程团队紧密合作
查看英文原文
Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.
As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch - and we're looking for the founding engineer to own the latter.
We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets.
The Opportunity
Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role.
Responsibilities
- Define Fuse's inference serving strategy and architecture from first principles.
- Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
- Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.
- Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
- Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
- Act as a direct technical owner of inference performance and reliability.
- Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
- Set the standards, tooling, and benchmarks this function will run on as it grows.
Requirements
· 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
- Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).
- Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster.
- Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
- A track record of making high-stakes architecture calls and owning the outcome.
- Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one.
Nice to Have
- Experience with Triton or custom ML inference/training frameworks.
- Experience with autoscaling or capacity planning for large-scale inference workloads.
- Exposure to multi-tenant serving or SLA-driven infrastructure.
- Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
- Familiarity with Kubernetes/Slurm for cluster orchestration.
- Interest or experience in energy markets, grid systems, or sustainability-focused compute.
Benefits
- Competitive salary and an equity sign-on bonus.
- Biannual bonus scheme.
- Fully expensed tech to match your needs.
- Breakfast and dinner allowance for office based employees.
Originally posted on Himalayas