人工智能基础设施/机器学习工程师
AI Infrastructure / ML Engineer
我们正在招聘一名AI基础设施/机器学习工程师,帮助优化[CapaCloud](https://capa.cloud/ "CapaCloud")以支持AI工作负载、模型部署、训练、推理和可扩展的GPU利用。
你将与基础设施工程师紧密合作,确保平台支持初创公司、研究人员和企业用户的现代AI工作流。
这个职位适合对AI系统、MLOps和大规模GPU计算充满热情的人。
## 主要职责
* 构建和优化AI部署流水线
* 提升AI应用的GPU工作负载效率
* 支持AI训练和推理基础设施
* 优化PyTorch、TensorFlow和LLM工作负载的性能
* 构建可扩展的API和推理系统
* 开发基准测试和性能测试工具
* 与基础设施团队协作开发编排系统
* 支持模型部署和容器化AI工作负载
* 提升AI用户的开发体验
* 监控和优化AI计算性能
## 必要技能与经验
* 具有AI/ML基础设施和MLOps经验
* 强大的Python编程能力
* 具有PyTorch、TensorFlow或JAX经验
* 具有GPU计算和CUDA环境经验
* 熟悉容器化部署系统
* 具有在生产环境中部署AI模型的经验
* 了解推理优化技术
* 具有API和后端系统经验
* 强大的调试和分析能力
## 优选条件
* 具有LLM基础设施经验
* 熟悉Hugging Face生态系统
* 具有分布式训练系统经验
* 了解Kubernetes和编排系统
* 具有AI推理优化工具经验
* 有开源AI贡献
## 成功的标准
* 高性能的AI部署基础设施
* AI工作负载的优化GPU利用率
* AI开发者的顺畅入职体验
* 可靠的推理和训练系统
* 跨工作负载的强基准性能
## 工作类型
* 全职
* 远程办公
查看英文原文
We are hiring an AI Infrastructure / ML Engineer to help optimize [CapaCloud ](https://capa.cloud/ "CapaCloud")for AI workloads, model deployment, training, inference, and scalable GPU utilization.
You will work closely with infrastructure engineers to ensure the platform supports modern AI workflows for startups, researchers, and enterprise users.
This role is ideal for someone passionate about AI systems, MLOps, and large-scale GPU computing.
## Key Responsibilities
* Build and optimize AI deployment pipelines
* Improve GPU workload efficiency for AI applications
* Support AI training and inference infrastructure
* Optimize performance for PyTorch, TensorFlow, and LLM workloads
* Build scalable APIs and inference systems
* Develop benchmarking and performance testing tools
* Collaborate with infrastructure teams on orchestration systems
* Support model deployment and containerized AI workloads
* Improve developer experience for AI users
* Monitor and optimize AI compute performance
## Required Skills & Experience
* Experience with AI/ML infrastructure and MLOps
* Strong Python programming skills
* Experience with PyTorch, TensorFlow, or JAX
* Experience with GPU computing and CUDA environments
* Familiarity with containerized deployment systems
* Experience deploying AI models in production
* Understanding of inference optimization techniques
* Experience with APIs and backend systems
* Strong debugging and analytical skills
## Nice To Have
* Experience with LLM infrastructure
* Familiarity with Hugging Face ecosystem
* Experience with distributed training systems
* Knowledge of Kubernetes and orchestration systems
* Experience with AI inference optimization tools
* Open-source AI contributions
## What Success Looks Like
* High-performance AI deployment infrastructure
* Optimized GPU utilization for AI workloads
* Smooth onboarding for AI developers
* Reliable inference and training systems
* Strong benchmark performance across workloads
## Employment Type
* Full-time
* Remote