工程技负责人(vNode)
Engineering Tech Lead (vNode)
作为vCluster Labs的工程技术负责人,你不仅仅是交付容器运行时功能;你正在定义Kubernetes操作员如何实现虚拟机级别的租户隔离,而无需承担虚拟机的代价。vNode用基于Linux用户命名空间和seccomp的运行时替代了虚拟kubelet和微虚拟机,而坐在这个职位上的人将决定该运行时的下一步发展方向。你将直接与vNode的创始工程师合作,为团队设定技术标准,并交付决定AI云和受监管企业是否将vNode作为默认隔离层的工作。
作为工程技术人员负责人,你的职责包括:
- 负责vNode的技术执行:推动vNode如何封装containerd、与kubelet集成以及暴露安全的隔离原语的架构设计。你将设定哪些内容可以发布、哪些需要推迟、哪些需要重新设计的标准。
- 深入研究容器运行时和隔离机制:领导vNode与containerd、Kata Containers、gVisor、runc和内核的对接工作。你将是能够解释(并改进)从Pod规范到在受限用户命名空间下运行进程之间具体发生了什么的人。
- 交付kubelet集成接口:负责vNode如何接入节点生命周期:CRI、kubelet设备插件、cgroups v2、驱逐机制,以及Kubernetes节点模型与不假设每个节点只有一个租户的运行时之间的边缘问题。
- 提升工程标准:进行技术设计评审,设定测试隔离保证的模式,并指导与你一起工作的工程师。你不是人事经理,但你是团队中其他人效仿的工程师。
- 成为vNode的首个用户:在客户看到之前,内部运行vNode的vCluster Platform租户集群。你将连接AI云运营商的需求与vNode在生产环境中的实际表现。
- 在外部代表vNode:在合适的地方贡献上游代码(如containerd、runc、Kubernetes SIG-Node),撰写解释基于命名空间的隔离为何是正确答案的技术文章,并在时机合适时在KubeCon级别的场合代表vCluster Labs发言。
如果你具备以下条件,这个职位可能适合你:
- 深厚的容器运行时经验:你有直接针对containerd的生产环境工作交付经验,而不仅仅通过Docker或Kubernetes使用它。对Kata Containers、gVisor或其他沙箱/隔离运行时有直接经验是加分项。
查看英文原文
As an Engineering Tech Lead at vCluster Labs, you aren't just shipping container runtime features; you are defining how Kubernetes operators get VM-grade tenant isolation without the VM tax. vNode replaces virtual kubelets and microVMs with a runtime built on Linux user namespaces and seccomp, and the person in this seat owns where that runtime goes next. You will partner directly with the vNode founding engineers, run the technical bar for the team, and ship the work that decides whether AI Clouds and regulated enterprises can adopt vNode as their default isolation layer.
As an Engineering Tech Lead, your role will include:
- Owning the vNode technical execution: Drive the architecture for how vNode wraps containerd, integrates with the kubelet, and exposes safe isolation primitives. You will set the bar for what ships, what gets deferred, and what gets redesigned.
- Going deep on container runtimes and isolation: Lead the work where vNode meets containerd, Kata Containers, gVisor, runc, and the kernel. You will be the person who can explain (and improve) exactly what happens between a Pod spec and a process running under a constrained user namespace with a tight seccomp profile.
- Shipping the kubelet integration surface: Own how vNode plugs into the node lifecycle: CRI, kubelet device plugins, cgroups v2, eviction, and the rough edges between Kubernetes' node model and a runtime that does not assume one tenant per node.
- Raising the engineering bar: Run technical design reviews, set the pattern for testing isolation guarantees, and mentor the engineers shipping alongside you. You are not a people manager, but you are the engineer the team copies.
- Being Customer Zero for vNode: Run vNode against vCluster Platform tenant clusters internally before customers see it. You will close the loop between what AI Cloud operators need and what vNode actually does in production.
- Representing vNode externally: Contribute upstream where it matters (containerd, runc, Kubernetes SIG-Node), write the technical posts that explain why namespace-based isolation is the right answer, and represent vCluster Labs at KubeCon-class venues when the timing is right.
This role could be a fit for you if you bring:
- Deep container runtime experience: You have shipped production work against containerd directly, not just consumed it through Docker or Kubernetes. Direct experience with Kata Containers, gVisor, or another sandboxed/isolated runtime is a strong plus.
- Kubernetes node-level depth: You have worked inside the kubelet, the CRI layer, or a node-resident agent. You know what cgroups v2, OCI hooks, and the kubelet's PLEG do and where they break.
- Go systems programming chops: You write production Go for systems-level code (syscalls, namespaces, file descriptors, process lifecycle), not just service handlers.
- Linux isolation fluency: User namespaces, seccomp-bpf, capabilities, and Landlock are not abstract concepts; you have shipped against them and can reason about their failure modes.
- Tech Lead instincts: You set technical direction by writing the design doc, prototyping the hard part, and then bringing the team along. You raise the bar without becoming the bottleneck.
Bonus points for:
- Upstream contribution history: Meaningful commits to containerd, runc, Kata, gVisor, Kubernetes SIG-Node, or related projects.
- Tenant Isolation domain expertise: You have built or operated infrastructure where the threat model includes hostile workloads on shared hosts (AI Cloud operators, multi-tenant SaaS, regulated industries).
- Public technical voice: Talks, posts, or RFCs that move the conversation on container isolation.
About vCluster Labs
We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.
We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.
We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.
Benefits
We offer the following benefits:
- Competitive Salary: We offer a competitive compensation package, including equity.
- Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).
- Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.
- Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.
Culture & Values
At vCluster Labs, we value and stand for:
- Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.
- Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.
- Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.
- Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy — the strongest ideas win, no matter who or where they come from.
- Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.
Compensation Range: $138K - $220K
Originally posted on Himalayas