高级AI DevOps / LLMOps
Senior AI DevOps / LLMOps
在 TechBiz Global,我们为组合中的顶级客户提供建设服务。我们正在寻找一位高级 AI DevOps / LLMOps 专家加入其中一家客户的团队。如果你正在寻找一个在创新环境中成长的激动人心的机会,这可能是适合你的选择
关键职责
· 构建到生产环境的自动化
- 设计并实现针对 AI 的强大 CI/CD 流水线,涵盖模型权重、数据集版本控制和应用代码。
- 开发专门的 PromptOps 工作流,确保系统提示被版本控制、进行回归测试,并以与传统代码相同的严谨性进行部署。
- 自动化代理工作流的部署,管理状态化 AI 交互和多代理交接的复杂性。
- 2. AI 基础设施即代码(IaC)
- 使用 Terraform、Pulumi 或 Ansible 配置和管理高性能计算环境(GPU 集群、TPU 泡泡)。
- 定义并执行 AI 端点的 Policy-as-Code,确保符合安全、成本使用限制和数据驻留要求。
- 在混合基础设施中保持一致的环境,确保本地开发与云生产之间的无缝一致性。
- 3. 安全实验与受控发布
- 架构 AI 的渐进式交付策略,包括金丝雀发布、蓝绿部署和影子(新模型与生产并行运行以比较输出)。
- 在流水线中构建“循环评估”门,自动在发布前测试偏差、幻觉和性能下降。
- 实现专为 LLM 输出和代理行为设计的 A/B 测试框架。
- 4. 监控与可观测性
- 建立对推理端点的深度可观测性,跟踪如每秒令牌数、延迟和模型准确性的漂移等指标。
- 集成反馈循环,捕捉生产中的“边缘案例”以反馈到训练和微调流水线。
- 要求
- 必须具备的技术技能:
- 编排:精通 Kubernetes(K8s),特别是 KubeFlow、Ray 或 NVIDIA Triton。
- CI/CD 与 IaC:精通 GitHub Actions/GitLab CI,以及 Terraform 或 Pulumi。
- AI 工具:有 Weights & Biases、MLflow、LangSmith 或 Arize Phoenix 的使用经验。
- 硬件:了解 GPU 虚拟化、CUDA 驱动程序和本地硬件管理。
- 安全:熟悉 Open Policy
查看英文原文
At TechBiz Global, we are providing recruitment service to our TOP clients from our portfolio. We are currently seeking an Senior AI DevOps / LLMOpsspecialist to join one of our clients' teams. If you're looking for an exciting opportunity to grow in a innovative environment, this could be the perfect fit for you.
Key Responsibilities
· Automation of Build-to-Production
- Design and implement robust CI/CD pipelines tailored for AI, covering model weights,
dataset versioning, and application code.
- Develop specialized workflows for PromptOps, ensuring that system prompts are
version-controlled, tested for regressions, and deployed with the same rigor as traditional
code.
- Automate the deployment of Agentic workflows, managing the complexities of stateful
AI interactions and multi-agent handoffs.
2. AI Infrastructure as Code (IaC)
- Provision and manage high-performance compute environments (GPU clusters, TPU
pods) using Terraform, Pulumi, or Ansible.
- Define and enforce Policy-as-Code for AI endpoints to ensure compliance with security,
cost-usage limits, and data residency requirements.
- Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless
parity between On-Premises development and Cloud production.
3. Safe Experimentation & Controlled Releases
- Architect Progressive Delivery strategies for AI, including Canary releases, Blue-Green
deployments, and Shadowing (where new models run in parallel with production to
compare outputs).
- Build “Evaluation-in-the-Loop” gates within the pipeline to automatically test for bias,
hallucination, and performance degradation before a release.
- Implement A/B testing frameworks specifically designed for LLM outputs and agentic
behavior.
4. Monitoring & Observability
- Establish deep observability into Inference Endpoints, tracking metrics like tokens-per-
second, latency, and drift in model accuracy.
- Integrate feedback loops that capture production “edge cases” to feed back into the
training and fine-tuning pipelines.
Requirements
Must-Have Technical Skills:
- Orchestration: Advanced Kubernetes (K8s) skills, specifically with KubeFlow, Ray, or
NVIDIA Triton.
- CI/CD & IaC: Expertise in GitHub Actions/GitLab CI, and Terraform or Pulumi.
- AI Tooling: Experience with Weights & Biases, MLflow, LangSmith, or Arize
Phoenix.
- Hardware: Understanding of GPU virtualization, CUDA drivers, and on-premises
hardware management.
- Security: Familiarity with Open Policy Agent (OPA) and secret management (Vault).
Experience:
- 10+ years in DevOps, SRE, or Cloud Engineering.
- 2+ years of hands-on experience in MLOps or LLMOps, specifically moving LLMs
from notebook to production.
- Proven experience managing Hybrid Cloud environments (e.g., AWS/Azure + Private
Data Center).
Highlights
- full time and remote job
- fluent English is needed
Originally posted on Himalayas