远程工作雷达

技术成员 - 强化学习环境

Member of Technical Staff - RL Environments

其他全球可投(据职位描述推断)
公司Cohere
薪资295,000 - 535,000 CAD
工作地点London
地域资格全球可投(据职位描述推断)
时区要求无特别要求
用工类型FullTime
发布时间29 天前
数据来源Ashby
前往企业招聘页投递 →
全球可投:该职位未限制候选人所在地区。仍需注意薪资可能按地区折算,以及实际签约方式(正式雇佣 / 独立合同)。

我们是谁?

Cohere 是领先的以安全为先的企业人工智能公司。我们构建前沿的基础 AI 模型和端到端产品,旨在解决现实世界中的业务问题。

我们正在为正在构建 AI 系统的企业训练和部署前沿模型。我们认为我们的工作对 AI 的广泛应用至关重要,我们正在寻找希望参与其中的人才。

我们对所构建的东西充满热情。我们每个人都负责提升模型的能力以及为客户创造的价值。Cohere 是一个由研究人员、工程师、设计师等组成的团队,大家都对自己的专业充满热情。

我们是一家总部位于多伦多的全球科技公司,在伦敦、纽约市、旧金山、蒙特利尔、巴黎、柏林和首尔设有主要办公室。加入我们吧!

职位概览:
构建能够协助企业任何工作的 AI 代理是一个具有挑战性且开放性的问题。解决这个问题的关键之一是尽可能真实地复制实际工作环境——用困难的任务来填充它们,创建合理的输入数据,并定义完成工作的明确奖励。我们构建了许多强化学习(RL)环境,然后将我们的代理放入其中进行评估或训练。

在这个职位中,你将负责创建这些 RL 环境,在其中运行 AI 代理,并在过程中改进代理和环境。结果会传递给客户,客户的反馈会再次输入——代理/环境改进循环将持续进行。

主要职责:
在这个领域有很多开放性问题。作为技术员工,RL 环境方向,你将:

- 构建针对不同代理能力和行业领域的 RL 环境

- 在这些环境中训练和评估代理

- 让所有组件协同工作:任务、数据、工具实现和验证器

- 在建模和产品之间协作,识别代理性能的差距,并改进代理和环境

- 与外部供应商合作,创建高质量、专家构建的 RL 环境,并构建工具以确保任务、数据和验证器的高质量

- 自动化发现模型能力差距,并在评估和训练期间系统地衡量代理性能

任职要求:

如果你符合以下条件,可能适合这个职位:

- 你曾构建过代理并针对特定行业使用场景进行优化

查看英文原文

Who are we?

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.

We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!

Role Overview:
Building AI agents that can assist with any kind of enterprise work is a challenging, open-ended problem. One key piece of solving it is replicating real work environments as realistically as possible - filling them with hard tasks to solve, creating plausible input data, and defining clear rewards for completing the work the right way. We build many of these reinforcement learning (RL) environments, then drop our agents into them to evaluate or train them.

In this role, you are responsible for creating these RL environments, running AI agents inside them, and improving both the agents and the environments in the process. The results reach customers, whose feedback feeds back in - and the agent/environment improvement loop continues.

Key Responsibilities:
There are many open problems in this space. As a Member of Technical Staff, RL Environments, you will:

- Build new RL environments targeting different agentic capabilities and industry areas

- Train and evaluate agents in those environments

- Make all the pieces work together: tasks, data, tool implementations, and verifiers

- Work across modeling and product to identify gaps in agent performance, and improve both the agents and the environments

- Work with external vendors to create high-quality, expert-built RL environments, and build tools to ensure high task, data, and verifier quality

- Automate the discovery of model capability gaps, and systematically measure agent performance during evals and training

Qualifications:

You may be a good fit if:

- You have engineered agents and optimized them for specific industry use cases

- You have spent dozens of hours reviewing agent trajectories to pinpoint exact failure points and fix them with model training or harness engineering

- You obsess over measuring agentic capabilities and turning that into a repeatable process

- You have had many debates about what a good outcome from an AI agent should look like, you translated that into verifier implementations and tuned the reward designs

- You have designed and run annotation workflows to surface insights into agent performance and verify data quality

- You have built synthetic data pipelines to scale eval and training efforts

- You use agents yourself in your daily work, and have stories about how you improved your setup to 10x your productivity

- A plus: you have experience with training with RL: scaling, troubleshooting and tuning the environments

If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply.

Working Location:
This role can be based remotely or from one of our office locations listed on the job description - there is no minimum in-office qualification requirement. We care most about hiring exceptional people regardless of locations, though please check the location listed on the posting for guidance around the core time zone or working hours alignment expected for the role.

FULL-TIME EMPLOYEES AT COHERE ENJOY THESE PERKS:

- A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.

- Full health and dental benefits, including a separate budget for mental health.

- RRSP matching, 401K, Pension Scheme.

- 100% Parental Leave top-up for up to 6 months, for either parent.

- Annual enrichment benefits:

Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.

Education & learning stipend for conferences, courses, and coaching.

- 6 weeks of paid vacation (30 working days!)

- Budget for traveling to other offices if you are remote, plus an annual company offsite.

HOW AND WHERE WE WORK:

- Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.

- For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.

- For those not near an office: a co-working benefit so you can work alongside others in your city.

- Everyone receives a $500 home office stipend to set up your workspace properly.

If any of the above doesn’t line up exactly with your experience, we still encourage you to apply.

We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form https://docs.google.com/forms/d/12a6IrLdF3kI2nonKSr4tiFuz18rLQbaeYV-JM9L4o9Q/edit, and we will work together to meet your needs.

We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.

Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers https://cohere.com/careers page.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

技术成员(多语言)

CohereLondon / New York / Paris / Toronto / Mont295,000 - 535,000 CADFullTime29 天前
其他全球可投(据职位描述推断)

← 返回全部职位