远程工作雷达

工程经理,运行时框架

Engineering Manager, Runtime Fabric

其他未标注地域
公司Baseten
薪资$165,000 - $330,000
工作地点San Francisco
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-06-09
数据来源Ashby
前往企业招聘页投递 →

关于Baseten

Baseten为全球最具活力的AI公司提供关键任务推理,如Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma和Writer。通过结合应用AI研究、灵活的基础设施和无缝的开发者工具,我们使处于AI前沿的公司能够将最先进的模型投入生产。我们正在快速成长,并最近完成了15亿美元的F轮融资,详情请见https://www.baseten.co/blog/announcing-our-series-f/,由Altimeter Capital、Conviction Partners和Spark Capital领投。加入我们,帮助构建工程师们用来发布AI产品的平台。

职位描述

容器运行时是为通用软件工作负载设计的。AI推理不是通用的工作负载。

在生产规模上运行大型模型会暴露容器堆栈每一层的缺陷:运行时不了解GPU内存限制,当模型需要扩展到数千个副本时,镜像需要数分钟才能拉取,以及没有为生产AI所需的多租户服务环境设计的隔离机制。行业过去十年依赖的工具并不是为此而构建的,仅在更高层次进行修补只能解决部分问题。

Baseten掌控整个流程,从开发者推送模型的那一刻起,到请求获得响应的那一刻为止。这种垂直掌控意味着我们可以从根源上解决这些问题。Runtime Fabric团队正是这样做的:专门为AI推理工作负载构建容器运行时和存储层,由一些全球顶尖的containerd维护者领导。

作为Runtime Fabric团队的工程经理,您将领导这项工作,设定技术方向,培养一支世界级的系统工程师团队,并确保团队的产出不仅塑造Baseten的基础设施,也影响整个开源容器生态系统。如果您曾参与过containerd、runc或相关OCI项目,并准备好带领一个团队解决当今最困难的基础设施问题,我们很期待与您交谈。

职责

团队领导与文化

- 招募、雇佣和发展一支具有深厚容器和Linux专业知识的系统工程师高绩效团队。

- 培养技术严谨性、开源贡献和持续改进的文化。

- 为您的直接下属提供定期指导、反馈和职业发展支持。

- 与工程领导合作,定义长期愿景和路线图

查看英文原文

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

Container runtimes were designed for general-purpose software workloads. AI inference is not a general-purpose workload.

Running large models at production scale exposes cracks in every layer of the container stack: runtimes unaware of GPU memory constraints, images that take minutes to pull when a model needs to scale to thousands of replicas, and isolation mechanisms that weren't designed for the multi-tenant serving environments that production AI requires. The tools the industry has relied on for a decade weren't built for this, and patching around those limitations at higher layers only goes so far.

Baseten owns the entire pipeline, from the moment a developer pushes a model to the moment a request gets a response. That vertical ownership means we can fix these problems at the root. The Runtime Fabric team is doing exactly that: purpose-building the container runtime and storage layers for AI inference workloads, led by some of the world's top containerd maintainers.

As Engineering Manager of the Runtime Fabric team, you will lead this work, setting technical direction, growing a world-class team of systems engineers, and ensuring the team's output shapes not just Baseten's infrastructure but the open-source container ecosystem at large. If you've contributed to containerd, runc, or related OCI projects and are ready to lead a team solving some of the hardest problems in infrastructure today, we'd love to talk.

RESPONSIBILITIES

Team Leadership & Culture

- Recruit, hire, and develop a high-performing team of systems engineers with deep container and Linux expertise.

- Foster a culture of technical rigor, open-source contribution, and continuous improvement.

- Provide regular coaching, feedback, and career development support to your direct reports.

- Partner with engineering leadership to define the long-term vision and roadmap for container runtime and storage infrastructure.

Technical Direction

- Guide the team in extending and hardening containerd, runc, and related OCI ecosystem projects to meet the GPU-specific requirements of production AI inference, including startup performance, GPU device access, and multi-tenant isolation.

- Oversee the architecture and evolution of the Baseten Delivery Network: the tiered caching and weight delivery system that makes cold starts 2–3x faster and eliminates thundering herd failures during burst scaling events.

- Drive the expansion of BDN's architecture, currently focused on model weights, to container images, training checkpoints, and deployment artifacts.

- Provide technical oversight on GPU-aware isolation mechanisms for multi-tenant inference, including secure container runtimes, Linux namespace hardening, and longer-term micro-VM integration.

- Ensure the team maintains end-to-end ownership of the container startup performance path, from snapshotter initialization through weight delivery to first inference request.

- Champion the team's contributions back to the open-source containerd ecosystem alongside a team of core maintainers.

Cross-Functional Partnership

- Act as the primary advocate for Runtime Fabric across the organization, ensuring upstream and downstream teams have the integration support they need.

- Collaborate with product and engineering stakeholders to prioritize investments based on business impact and infrastructure reliability.

- Communicate team progress, technical trade-offs, and architectural decisions clearly to leadership.

REQUIREMENTS

- Proven experience managing and growing engineering teams in a systems, infrastructure, or low-level runtime context.

- Deep familiarity with the Linux container ecosystem: containerd, runc, OCI Runtime Spec, Linux namespaces, and cgroups, with the ability to engage credibly in code reviews and architectural discussions.

- Contributions to containerd/containerd, opencontainers/runc, google/gvisor, kata-containers/kata-containers, or closely related open-source projects.

- Strong systems programming background in Go and/or C/C++.

- Experience with distributed storage systems, content-addressable storage, or large-scale caching infrastructure.

- Understanding of how container images are structured, stored, and delivered at scale.

- Strong written and verbal communication skills, with the ability to influence without authority across teams.

NICE TO HAVE

- Experience with GPU device access in containers: NVIDIA Container Toolkit, CDI (Container Device Interface), or GPU-aware scheduling.

- Familiarity with lazy-loading snapshotters (stargz, soci, EROFS/Nydus) or peer-to-peer image distribution.

- Experience with secure container runtimes (gVisor, Sysbox) or micro-VM technologies (Firecracker, Cloud Hypervisor).

- Understanding of containerd's shim API (v2) and experience building custom shim implementations.

- Background in multi-tenant infrastructure or security-sensitive serving environments.

BENEFITS

- Competitive compensation, including meaningful equity

- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents

- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

- Paid parental leave

- Fertility and family-building stipend through Carrot

- Company-facilitated 401(k)

- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

AI推理工程师

BasetenSan Francisco / Remote / Toronto / New Yor$165,000 - $330,000FullTime2026-08-03
AI开发工程全球可投

← 返回全部职位