远程工作雷达

研究工程师 (强化学习)

Research Engineer (Reinforcement Learning)

开发工程全球可投
公司LiveKit
薪资未公开
工作地点不限地点
地域资格全球可投
时区要求无特别要求
用工类型未标注
发布时间28 天前
数据来源Real Work From Anywhere
前往 Real Work From Anywhere 查看并投递 →
全球可投:该职位未限制候选人所在地区。仍需注意薪资可能按地区折算,以及实际签约方式(正式雇佣 / 独立合同)。

LiveKit 正在构建面向语音驱动计算时代的基础架构层。我们的平台为开发者提供了一切必要的工具,以构建、测试、部署、扩展和监控生产环境中的智能体。LiveKit 成立于 2021 年,为 OpenAI、xAI、Salesforce、Coursera、Spotify 以及数千家其他公司提供语音 AI 应用,每年共同处理数十亿次通话。
关于该职位
我们正在寻找一位杰出的工程师来在 LiveKit 构建后训练(post-training)能力。我们的智能体通过语音运行,越来越多地通过 SMS 和聊天等文本渠道运行,而有趣的问题出现在长期范围内:在多个会话中保持有用性,在长时间内处理累积的上下文,并在实时对话中可靠地使用工具。
你将负责
构建我们的模型训练所针对的环境和验证器
负责合成数据管道,从生成到质量检查
端到端运行训练实验,并解释模型的变化
构建发布必须通过的评估
选择并适应适用于我们任务的开源权重基础模型
确保训练行为对语音和文本智能体都有效
将模型部署到生产环境,并根据真实使用情况持续改进
你应具备
扎实的 Python 工程师技能
能够将模型从原始数据带到生产环境
将数据视为产品:覆盖率、多样性、泄露
假设模型会利用弱奖励,并设计应对方案
熟悉 GPU 并诚实面对其局限性
知道何时需要训练,何时不需要
在远程环境中舒适地协作
加分项
具有后训练经验:微调、奖励设计或强化学习,如 GRP ORL 或微调框架如 TRL、verl 或 OpenRLHF,或自己编写的训练循环
使用 vLLM 或 SGLang 实现快速部署,使用 FSDP 进行多 GPU 训练
使用训练工具或多轮智能体
执行沙盒、验证器、评估套件或其他工程师依赖的工具
开源权重系列如 Qwen 或 Llama,LoRA 及类似技术
我们对你的承诺
塑造一个快速增长的开发者平台品牌的机会
与一支小型、资深团队合作,他们高度重视工艺和创造力
有竞争力的薪资和股权包
健康、牙科和视力福利
灵活的休假政策
LiveKit 是一家平等机会雇主,不基于任何受适用法律保护的特征进行歧视。如果你在申请或面试过程中需要合理的便利措施,请告知我们。

查看英文原文

About LiveKit LiveKit is building the infrastructure layer for the voice-driven era of computing. Our platform gives developers everything they need to build, test, deploy, scale, and observe agents in production. Founded in 2021, LiveKit powers voice AI applications for OpenAI, xAI, Salesforce, Coursera, Spotify, and thousands of others, collectively facilitating billions of calls each year.About This RoleWe are looking for an exceptional engineer to build post-training at LiveKit. Our agents run over voice and increasingly over text channels like SMS and chat, and the interesting problems show up over long horizons: staying useful across many sessions, working with context that accumulates over time, and using tools reliably in the middle of a live conversation.What You'll DoBuild the environments and verifiers our models train againstOwn the synthetic data pipeline, from generation through the quality gatesRun training experiments end to end, and explain what moved the modelBuild the evaluations a release has to clearChoose and adapt open-weight base models for our tasksMake trained behavior hold up for voice and text agents alikeShip models into production and keep improving them on real usageWho You AreA strong Python engineerHave carried a model from raw data through to productionTreat data as the product: coverage, diversity, leakageAssume a model will exploit a weak reward, and design against itComfortable with GPUs and honest about their limitsKnow when to train, and when not toComfortable working collaboratively in a remote environmentNice to HaveExperience with post-training: fine-tuning, reward design, or reinforcement learning such as GRPORL and fine-tuning frameworks such as TRL, verl, or OpenRLHF, or a training loop you wrote yourselfFast rollouts with vLLM or SGLang, multi-GPU training with FSDPTraining tool-using or multi-turn agentsExecution sandboxes, verifiers, eval harnesses, or tooling other engineers depend onOpen-weight families such as Qwen or Llama, LoRA and similarOur Commitment to YouThe opportunity to shape the brand of a fast-growing developer platformCollaboration with a small, senior team that deeply values craft and creativityCompetitive salary and equity packageHealth, dental, and vision benefitsFlexible vacation policyLiveKit is an equal opportunity employer and does not discriminate on the basis of any characteristic protected by applicable law. If you require a reasonable accommodation during the application or interview process, please contact recruiting@livekit.io.

本页面信息整理自 Real Work From Anywhere,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位