人工智能基础设施工程师
AI Infrastructure Engineer
作为vCluster的AI基础设施专家,你将直接与客户在他们旅程的最早和最关键阶段合作:从裸金属GPU节点到生产就绪的部署。这不是一个传统的专业服务角色;你在售前阶段作为价值验证项目的一部分,目标是实现生产环境。你将是新云或AI工厂在技术深度上接触的第一批团队成员之一,你制定的指南将为后续的招聘和客户扩展提供支持。
vCluster正在GPU AI云和构建AI工厂的企业中迅速获得认可:这些组织需要在裸金属GPU基础设施上提供Kubernetes作为托管服务,并且需要快速实现。这个职位的存在就是为了实现这一目标。
作为AI基础设施工程师,你的职责包括:
- 领导技术部署:为GPU neocloud和AI工厂客户提供端到端的技术部署,从初始的裸金属配置到验证过的vCluster环境。
- 基础设施优化:配置和排查裸金属GPU节点基础设施,包括CNI配置、GPU Operator设置、分布式存储后端以及RDMA/InfiniBand。
- 验证:部署并验证Kubernetes和vCluster,以提供基于GPU的托管K8s。
- 知识转移:与客户团队协作,建立自主性,确保他们能够独立操作和扩展平台。
- 通过文档进行扩展:记录可重复使用的指南和部署架构,使你的经验成为下一位客户的起点。
- 反馈循环:与工程和产品团队合作,揭示常见的基础设施挑战,作为现场到产品路线图的直接反馈渠道。
- 战略合作:在需要深入基础设施工作的预销售过程中,与销售团队合作,实现有意义的价值验证。
如果你具备以下条件,这个职位可能适合你:
- 精通生产环境K8s:5年以上在生产环境中部署和运维Kubernetes的经验,最好是在裸金属或高复杂度环境中。
- GPU熟练:对NVIDIA GPU Operators、CUDA工具链和GPU节点系统级配置有实际了解。
- 网络基础:对CNI插件、覆盖网络、负载均衡和分层环境中的连接诊断有深入理解。
- 存储专长:有持久卷配置、CSI驱动程序和存储解决方案的经验。
查看英文原文
As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engagement scoped to reach production. You will be one of the first team members a neocloud or AI Factory engages with at a technical depth, and the playbooks you develop will scale the motion for the next hire and customer.
vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen.
As an AI Infrastructure Engineer, your role will include:
- Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment.
- Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand.
- Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.
- Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently.
- Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start.
- Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap.
- Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.
This role could be a fit for you if you bring:
- Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.
- GPU Fluency: Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.
- Networking Fundamentals: Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments.
- Storage Expertise: Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn.
- Operational Agility: Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time.
- Modern Tech Mindset: You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services.
BONUS POINTS FOR:
- Automation Skills: Experience writing automation scripts with Bash, Python, or Go.
- Kubernetes Depth: Relevant certifications such as CKA (Certified Kubernetes Administrator) or experience writing Kubernetes Operators.
- AI/ML Familiarity: Experience with inference serving, GPU scheduling, and the tooling around LLM deployment.
- Documentation: Experience building AI Automation in documentation to contribute to a shared knowledge base.
ABOUT VCLUSTER LABS
We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.
We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.
We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.
Benefits
We offer the following benefits:
- Competitive Salary: We offer a competitive compensation package, including equity.
- Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).
- Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.
- Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.
CULTURE & VALUES
At vCluster Labs, we value and stand for:
1. Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.
2. Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.
3. Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.
4. Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy — the strongest ideas win, no matter who or where they come from.
5. Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.