远程工作雷达

软件工程师 - GPU 网络与分布式系统

Software Engineer - GPU Networking & Distributed Systems

开发工程未标注地域
公司Baseten
薪资$165,000 - $330,000
工作地点San Francisco / Toronto / New York / Montreal
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-02-23
数据来源Ashby
前往企业招聘页投递 →

ABOUT BASETEN

Baseten 为全球最具活力的 AI 公司提供关键任务的推理支持,如 Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma 和 Writer。通过结合应用 AI 研究、灵活的基础架构和无缝的开发者工具,我们使处于 AI 前沿的公司能够将前沿模型投入生产。我们正在快速成长,并最近完成了 1.5 亿美元的 F 轮融资 https://www.baseten.co/blog/announcing-our-series-f/,由 Altimeter Capital、Conviction Partners 和 Spark Capital 领投。加入我们,帮助构建工程师们用来部署 AI 产品的平台。

THE ROLE

在 Baseten,我们正在打造分布式异构 AI 硬件的全球操作系统。我们认为,随着 LLM 和多模态工作负载的扩展,网络就是计算机。我们正在寻找基础工程师来领导我们的 GPU 网络工作,使 RDMA 成为我们基础架构中的第一优先级构建模块,并解锁下一代分布式推理优化。

网络和计算不再属于不同的学科;它们正在融合。H100、B200 和 NVL72 架构的巨大吞吐量带来了新的方法,其中通信与计算一同进行优化。我们正进入一个网络成为主动加速器的时代,利用智能硬件卸载和直接互联,确保数据传输以线速运行。

在这个职位中,你将超越网络配置,设计统一数千个 GPU 的软件架构,形成一个连贯的操作系统。虽然你会利用开源生态系统的最佳实践,但不会受其限制。当现成的解决方案无法满足需求时,你将从零开始构建,工程化所需的原始功能,以协同优化通信和计算,用于解耦服务、宽专家并行(WideEP)以及加速冷启动。

RESPONSIBILITIES

- 让 RDMA 成为第一优先级:你将致力于将 RDMA/RoCE/InfiniBand 功能直接集成到我们的推理堆栈中,帮助我们超越 TCP/IP,实现带宽和延迟的量级提升。

- 优化分布式推理:你将实现并调优用于高效解耦 KV 缓存卸载和 WideEP 的网络层,确保我们的 MoE 模型在 NVLink 和 InfiniBand 上实现无缝通信。

- 为 LLM 实现无服务器级别的启动速度:你将深入研究检查点和存储机制,以实现低延迟的模型加载和快速启动。

查看英文原文

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations.

Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed.

In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts.

RESPONSIBILITIES

- Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack, helping us move beyond TCP/IP to unlock order-of-magnitude improvements in bandwidth and latency.

- Optimize Distributed Inference: You will implement and tune the networking layers necessary for efficient Disaggregated KV Cache Offload and WideEP, ensuring seamless communication across NVLink and InfiniBand for our MoE models.

- Enable Serverless-Grade Startup Speeds for LLMs: You will work deeply with checkpointing and storage mechanisms to enable sub-10-second startup for trillion-parameter models.

- Deep-Dive into Hardware: You will characterize and validate networking performance on bleeding-edge clusters (H100/H200, B200/B300, GB200/300 NVL72), writing the acceptance tests that ensure our hardware delivers peak achievable throughput and minimal latency.

- Build Observability: You will design the tools that let us visualize packet flow, congestion, and effective bandwidth across the GPU interconnects, helping us diagnose complex distributed system behaviors.

- Optimize Kernels: You will work with communication libraries (NCCL, NVSHMEM) and potentially write custom communication kernels to overlap compute and data transfer.

REQUIREMENTS

- You have deep experience with high-performance networking protocols (InfiniBand, RoCE v2) and understand the physics of data movement.

- You are fluent in C++ or Python, with the ability to bridge the gap between high-level logic and hardware. You have a deep understanding of the memory hierarchy in modern NVIDIA architectures (H100/Blackwell) and know how to optimize for it.

- You like going deep. You aren't afraid to dive into TensorRT-LLM source code, write custom C++ / Python bindings, or debug NVLink topology issues.

- You know when to use an off-the-shelf solution and when we need to build a custom solution because the upstream tools (like standard Kubernetes networking) are too slow for our needs.

NICE TO HAVE

- Deep knowledge of NCCL, NVSHMEM, and UCX.

- Experience with Rust for systems-level or performance-critical networking code is a strong plus

- Experience with GPUDirect Storage (GDS) or high-performance filesystems like Weka or 3FS.

- Familiarity with TensorRT-LLM, vLLM, or SGLang.

- Experience running low-level benchmarks to "qualify" new hardware clusters.

THE TEAM

- Bleeding Edge Hardware: We are preparing to bring Blackwell (B200/B300) and then Rubin architectures online. You will be one of the first engineers in the industry optimizing networking for NVL72/GB300 racks.

- We go deep: We operate at every depth. Whether it’s tuning hardware interconnects, writing custom communication kernels, or designing distributed inference strategies, we work across the entire stack to deliver performance that goes far and beyond.

- High Impact: The networking optimizations you build will directly enable features that no one else in the industry has fully mastered yet, like seamless multi-node WideEP and instant model hydration.

BENEFITS

- Competitive compensation, including meaningful equity

- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents

- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

- Paid parental leave

- Fertility and family-building stipend through Carrot

- Company-facilitated 401(k)

- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

AI推理工程师

BasetenSan Francisco / Remote / Toronto / New Yor$165,000 - $330,000FullTime2026-08-03
AI开发工程全球可投

← 返回全部职位