远程工作雷达

首席机器学习运维工程师

Principal ML Ops Engineer

开发工程未标注地域
公司pragmatike
薪资未公开
工作地点Cambridge / New York / Chicago / San Francisco / Cambridge / Pennsylvania / New Jersey
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-04-22
数据来源Ashby
前往企业招聘页投递 →

位置:马萨诸塞州剑桥市(东部时间 / UTC -4)提供搬迁补贴或为非本地申请者提供远程选项
开始日期:尽快
语言:英语(必需)

关于该职位

Pragmatike 正代表一家由 MIT CSAIL 研究人员创立的快速发展的 AI 初创公司招聘,该公司被 GTM Capital 评为 Top 10 GenAI 公司。

我们正在寻找一位资深/首席 ML Ops 工程师,负责设计、实现和扩展公司的 ML 基础设施和生产级 AI 系统。这是一个高影响力、定义架构的角色,你将参与整个模型生命周期——训练、评估、部署、可观测性以及持续优化。

你将与 AI 研究人员、GPU 系统工程师、后端团队和产品利益相关者紧密合作,确保公司的大规模 AI 系统具备稳健性、高效性、自动化和生产级质量。这个职位适合已经大规模构建并主导过 ML 平台的人,能够推动战略规划及实际执行。

你将负责

- 设计、构建和扩展端到端的 ML Ops 流水线,包括训练、微调、评估、发布和监控。

- 设计可靠的模型部署、版本控制、可复现性和跨云和本地 GPU 集群的编排基础设施。

- 优化分布式系统中的计算资源使用(Kubernetes、自动扩展、缓存、GPU 分配、检查点工作流)。

- 主导 ML 系统的可观测性实现(监控漂移、性能、吞吐量、可靠性、成本)。

- 构建数据集整理、标注、特征流水线、评估以及 ML 模型的 CI/CD 自动化流程。

- 与研究人员合作,将模型投入生产,并加速训练/推理流水线。

- 建立 ML Ops 最佳实践、内部标准和跨团队工具。

- 指导工程师并影响整个 AI 平台的架构方向。

我们寻找

- 在大规模生产 ML 系统方面有深入的实际经验(期望具备资深/首席级别能力)。

- 在 ML Ops、分布式系统和云基础设施(AWS、GCP 或 Azure)方面有扎实背景。

- 精通 Python,并熟悉 TypeScript 或 Go 用于平台集成。

- 精通 ML 框架:PyTorch、Transformers、vLLM、Llama-factory、Megatron-LM、CUDA/GPU 加速(实际理解)

- 在容器化和编排方面有丰富经验

查看英文原文

Location: Cambridge, MA (Eastern Time / UTC -4) Relocation package available or Remote option for Out-Of-State applicants
Start date: ASAP
Languages: English (required)

ABOUT THE ROLE

Pragmatike is hiring on behalf of a fast-growing AI startup recognized as a Top 10 GenAI company by GTM Capital, founded by MIT CSAIL researchers.

We are seeking a Staff / Principal ML Ops Engineer to lead the design, implementation, and scaling of the companys ML infrastructure and production AI systems. This is a high-impact, architecture-defining role where youll work across the entire model lifecycletraining, evaluation, deployment, observability, and continuous optimization.

You will partner closely with AI researchers, GPU systems engineers, backend teams, and product stakeholders to ensure the companys large-scale AI systems are robust, efficient, automated, and production-grade. This role is ideal for someone who has already built and owned ML platforms at scale and can drive strategy as well as hands-on execution.

WHAT YOULL DO

- Architect, build, and scale the end-to-end ML Ops pipeline, including training, fine-tuning, evaluation, rollout, and monitoring.

- Design reliable infrastructure for model deployment, versioning, reproducibility, and orchestration across cloud and on-prem GPU clusters.

- Optimize compute usage across distributed systems (Kubernetes, autoscaling, caching, GPU allocation, checkpointing workflows).

- Lead the implementation of observability for ML systems (monitor drift, performance, throughput, reliability, cost).

- Build automated workflows for dataset curation, labeling, feature pipelines, evaluation, and CI/CD for ML models.

- Collaborate with researchers to productionize models and accelerate training/inference pipelines.

- Establish ML Ops best practices, internal standards, and cross-team tooling.

- Mentor engineers and influence architectural direction across the entire AI platform.

WHAT ARE LOOKING FOR

- Deep hands-on experience designing and operating production ML systems at scale (Staff/Principal-level expected).

- Strong background in ML Ops, distributed systems, and cloud infrastructure (AWS, GCP, or Azure).

- Proficiency with Python and familiarity with TypeScript or Go for platform integration.

- Expertise in ML frameworks: PyTorch, Transformers, vLLM, Llama-factory, Megatron-LM, CUDA / GPU acceleration (practical understanding)

- Strong experience with containerization and orchestration (Docker, Kubernetes, Helm, autoscaling).

- Deep understanding of ML lifecycle workflows: training, fine-tuning, evaluation, inference, model registries.

- Ability to lead technical strategy, collaborate cross-functionally, and operate in fast-paced environments

BONUS POINTS

- Experience deploying and operating LLMs and generative models in production at enterprise scale.

- Familiarity with DevOps, CI/CD, automated deployment pipelines, and infrastructure-as-code.

- Experience optimizing GPU clusters, scheduling, and distributed training frameworks.

- Prior startup experience or comfort operating with ambiguity and high ownership.

- Experience working with data engineering, feature pipelines, or real-time ML systems.

WHY THIS ROLE WILL PIVOT YOUR CAREER

- Research pedigree: MIT CSAIL founders recognized for breakthrough AI and systems contributions.

- Customer impact: Deploy AI solutions powering Fortune 500 clients.

- Industry momentum: Lab alumni have led high-value acquisitions (MosaicML Databricks, Run:AI Nvidia, W&B CoreWeave).

- Funding & growth: Oversubscribed seed round, next funding in 2026.

- Career growth & influence: Lead AI initiatives, optimize pipelines, and directly impact production AI systems at scale.

- Culture & autonomy: Own critical systems while collaborating with world-class engineers.

- Aspirational impact: Solve AI performance challenges few engineers ever face.

BENEFITS

- Competitive salary & equity options

- Sign-on bonus

- Health, Dental, and Vision

- 401k

Pragmatike is an Equal Opportunity Employer and is committed to providing equal employment opportunities to all applicants without discrimination. We recruit on behalf of our clients and prohibit discrimination and harassment based on race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training.We are committed to a fair and inclusive hiring process. We process your personal data solely for recruitment purposes, in accordance with applicable privacy laws, and maintain reasonable safeguards to protect your information. Your data may be shared with our client(s) for hiring consideration, but will not be disclosed to third parties outside of the recruitment process.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位