远程工作雷达

高级推理工程师

Senior Inference Engineer

AI开发工程全球可投
公司vclusterlabs
薪资$190,000 - $220,000
工作地点Remote / Canada / EMEA / United States
地域资格全球可投
时区要求无特别要求
用工类型FullTime
发布时间4 天前
数据来源Ashby
前往企业招聘页投递 →
全球可投:该职位未限制候选人所在地区。仍需注意薪资可能按地区折算,以及实际签约方式(正式雇佣 / 独立合同)。

作为vCluster Labs的高级推理工程师,你是我们首批招聘的工程师之一,负责推理系统。你将直接与首席技术官合作,从零开始构建平台的推理层,将模型转化为可扩展的、生产级的查询到响应流程。在此基础上,你将帮助引领vCluster在推理方面的工程方向,与产品团队合作,随着领域的发展共同决定下一步要构建的内容。

作为高级推理工程师,你的职责包括:

- 部署模型到生产环境:将大语言模型部署到GPU基础设施上的一个或多个机器上,从客户查询到响应的整个流程均由你负责。

- 服务框架:使用vLLM、SGLang或TensorRT-LLM搭建并运营服务基础设施。

- 优化可扩展性:应用量化、批处理、缓存和路由技术,在流量增长时控制延迟和成本。

- 编程,不只是配置:用Python或Golang构建真实的基础设施——这是一个工程岗位,而非研究或数据科学岗位。

- 主导路线图:与我们的首席技术官一起构建第一个版本,之后主导推理平台的发展,并与产品团队合作决定我们接下来要构建的内容。

如果你具备以下条件,这个职位可能适合你:

- 具有生产环境大语言模型服务经验:你曾使用vLLM、SGLang或TensorRT-LLM部署和提供大语言模型服务,最好是在以推理为核心的公司工作过。

- 具备推理优化的知识和技能:有量化、批处理、缓存和路由的实际操作经验,而不仅仅是了解这些术语。

- 具有实际编程经验:在Python或Golang方面有扎实的工程能力,并有实际的生产代码经验。

- 沟通能力:具备良好的沟通能力,能够向工程师和非技术利益相关者清晰地解释技术概念。

加分项:

- 熟悉容器化环境(Docker、Kubernetes)

- 具有常见机器学习框架(PyTorch、Transformers)的生成式AI实际经验

- 对GPU堆栈有良好理解:CUDA、NCCL、驱动程序及相关库

- 了解模型架构和微调方法

- 具有NVIDIA Dynamo使用经验

关于vCluster Labs

我们是全球AI基础设施领域的首选平台,被世界上发展最快的AI云构建者所信赖。我们是一家获得风险投资的初创公司,已从包括Khosla Ventures(OpenAI的早期投资者)在内的顶级投资者处筹集了超过2800万美元的资金。

查看英文原文

As a Senior Inference Engineer at vCluster Labs, you are the first engineer we're hiring to own inference. You'll partner directly with our CTO to build the platform's inference layer from the ground up, taking models and turning them into a production-grade, query-to-response pipeline running at scale. From there, you will help lead the engineering direction of inference at vCluster, partnering with Product to shape what we build next as the space evolves.

As a Senior Inference Engineer, your role will include:

- Deploying models to production: Take LLMs and put them into production across one or more machines on GPU infrastructure, owning the full pipeline from a customer's query to the served response.

- Serving frameworks: Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.

- Optimizing for scale: Apply quantization, batching, caching, and routing to keep latency and cost in check as traffic grows.

- Programming, not just configuring: Build real infrastructure in Python or Golang — this is an engineering role, not a research or data-science one.

- Owning the roadmap: Build the first iteration alongside our CTO, then take the lead on the inference platform and partner with Product to decide what we build next.

This role could be a fit for you if you bring:

- Production LLM serving experience: You've deployed and served LLMs using vLLM, SGLang, or TensorRT-LLM, ideally at a company built around inference at scale.

- Inference optimization know-how: Hands-on experience with quantization, batching, caching, and routing, not just familiarity with the terms.

- Hands-on programming experience: Strong engineering skills in Python or Golang, with real production code experience.

- Communication: Strong communication skills, explaining technical concepts clearly to both engineers and non-technical stakeholders.

Bonus points for:

- Familiarity with containerized environments (Docker, Kubernetes)

- Hands-on generative AI experience with common ML frameworks (PyTorch, Transformers)

- Good understanding of the GPU stack: CUDA, NCCL, drivers, and related libraries

- Knowledge of model architectures and fine-tuning approaches

- Experience with NVIDIA Dynamo

ABOUT VCLUSTER LABS

We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.

We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.

We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

Benefits

We offer the following benefits:

- Competitive Salary: We offer a competitive compensation package, including equity.

- Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).

- Flexible Working Schedule:  You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

- Workplace Flexibility:  We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

CULTURE & VALUES

At vCluster Labs, we value and stand for:

1. Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.

2. Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.

3. Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.

4. Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy — the strongest ideas win, no matter who or where they come from.

5. Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

区域销售总监,西部

vclusterlabsUSA - San Francisco / USA - Los Angeles$400,000 - $440,000FullTime昨天
市场运营全球可投(据职位描述推断)

社交媒体经理

vclusterlabsUnited States$105,000 - $125,000FullTime7 天前
市场运营全球可投(据职位描述推断)

高级产品经理

vclusterlabsRemote / Canada / EMEA / United States$170,000 - $200,000FullTime15 天前
职能支持全球可投

资深产品经理 (vMetal)

vclusterlabsRemote / Canada / EMEA / United States$170,000 - $210,000FullTime15 天前
职能支持全球可投

← 返回全部职位