远程工作雷达

高级机器学习系统工程师,框架与工具

Senior ML Systems Engineer, Frameworks & Tooling

开发工程未标注地域
公司Cohere
薪资295,000 - 535,000 CAD
工作地点London / San Francisco / New York / Paris / Toronto / Montreal
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2025-12-01
数据来源Ashby
前往企业招聘页投递 →

我们是谁?

Cohere 是一家以安全为先的企业人工智能公司,处于行业领先地位。我们构建前沿的基础 AI 模型和端到端产品,旨在解决现实世界中的业务问题。

我们正在为正在构建 AI 系统的企业训练和部署前沿模型。我们认为我们的工作对 AI 的广泛应用至关重要,我们正在寻找希望成为其中一员的人才。

我们对所构建的产品精益求精。我们每个人都负责提升模型的能力以及为客户创造的价值。Cohere 是一个由研究人员、工程师、设计师等组成的团队,他们对各自的专业充满热情。

我们是一家总部位于多伦多的全球科技公司,在伦敦、纽约市、旧金山、蒙特利尔、巴黎、柏林和首尔设有重要办事处。加入我们!

职位概览:

我们正在寻找一位高级工程师,帮助构建、维护和演进支持我们前沿规模语言模型的训练框架。这个职位位于大规模训练、分布式系统和 HPC 基础设施的交汇点。你将设计并维护使模型训练快速、可靠且可扩展的核心组件,并构建连接研究想法与数千块 GPU 的工具。

如果你喜欢在 ML 系统的全栈中工作,这个职位将为你提供机会和自主权,通过参与以下项目产生巨大影响:

- 构建高性能的数据加载和缓存管道。

- 在 ML 系统堆栈中实现性能分析。

- 开发内部指标和训练运行的监控功能。

- 构建可重复性和回归测试基础设施。

- 开发高性能、容错的分布式检查点系统。

主要职责:

- 构建并负责大型 LLM 训练的训练框架。

- 设计分布式训练抽象(数据/张量/流水线并行,FSDP/ZeRO 策略,内存管理,检查点)。

- 提升多节点集群(如 GB200/300、AMD、H200/100)上的训练吞吐量和稳定性。

- 开发和维护用于监控、日志记录、调试和开发者体验的工具。

- 与基础设施团队紧密合作,确保我们的集群、容器环境和硬件配置支持高性能训练。

- 调查并解决 ML 系统堆栈中的性能瓶颈。

- 构建可靠的分布式训练系统。

查看英文原文

Who are we?

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.

We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!

Role Overview:

We’re looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure. You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs.

If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact by working on projects such as:

- Building a high-performance data loading and caching pipeline.

- Implementing performance profiling across the ML systems stack

- Developing internal metrics and monitoring for training runs.

- Building reproducibility and regression testing infrastructure.

- Developing a performant fault-tolerant distributed checkpointing system.

Key Responsibilities:

- Build and own the training framework responsible for large-scale LLM training.

- Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing).

- Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100).

- Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.

- Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training.

- Investigate and resolve performance bottlenecks across the ML systems stack.

- Build robust systems that ensure reproducible, debuggable, large-scale runs.

Qualifications:

- Strong engineering experience in large-scale distributed training or HPC systems.
Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.

- Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).

- Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.

- Experience working with containerized environments (Docker, Singularity/Apptainer).

- A track record of building tools that increase developer velocity for ML teams.

- Excellent judgment around trade-offs: performance vs complexity, research velocity vs maintainability.

- Strong collaboration skills — you’ll work closely with infra, research, and deployment teams.

Any of the following would also be good to have for this role:

- Experience with training LLMs or other large transformer architectures.

- Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).

- Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches).

- Experience with data pipeline optimization, sharded datasets, or caching strategies.

- Background in performance engineering, profiling, or low-level systems.

Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP).

Working Location:
This role can be based remotely or from one of our office locations listed on the job description - there is no minimum in-office qualification requirement. We care most about hiring exceptional people regardless of locations, though please check the location listed on the posting for guidance around the core time zone or working hours alignment expected for the role.

FULL-TIME EMPLOYEES AT COHERE ENJOY THESE PERKS:

- A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.

- Full health and dental benefits, including a separate budget for mental health.

- RRSP matching, 401K, Pension Scheme.

- 100% Parental Leave top-up for up to 6 months, for either parent.

- Annual enrichment benefits:

Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.

Education & learning stipend for conferences, courses, and coaching.

- 6 weeks of paid vacation (30 working days!)

- Budget for traveling to other offices if you are remote, plus an annual company offsite.

HOW AND WHERE WE WORK:

- Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.

- For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.

- For those not near an office: a co-working benefit so you can work alongside others in your city.

- Everyone receives a $500 home office stipend to set up your workspace properly.

If any of the above doesn’t line up exactly with your experience, we still encourage you to apply.

We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form https://docs.google.com/forms/d/12a6IrLdF3kI2nonKSr4tiFuz18rLQbaeYV-JM9L4o9Q/edit, and we will work together to meet your needs.

We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.

Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers https://cohere.com/careers page.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位