远程工作雷达

技术成员(数据科学家,评估)

Member of Technical Staff (Data Scientist, Evals)

AI未标注地域
公司Perplexity
薪资$200,000 - $300,000
工作地点San Francisco / Palo Alto
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-06-29
数据来源Ashby
前往企业招聘页投递 →

Perplexity 通过以 LLM 为主导的搜索引擎和我们专业的数据源,每天为数千万用户提供可靠、高质量的答案。我们希望在最新模型发布时立即使用,但智能边界是不规则的,流行的基准测试并不能有效覆盖我们的使用场景。在这个职位中,你将构建专门的评估方案,以提升 Perplexity 的答案质量,涵盖基于搜索的 LLM 答案和其他用户常用场景。

职责

- 架构并维护自动化评估流程,以评估 Perplexity 产品中的答案质量,确保准确性和帮助性达到高标准

- 设计专门的评估集和方法,以衡量工具调用(尤其是网络搜索检索)对最终答案质量的影响

- 开发基于视觉语言模型(VLM)的解决方案,以程序化方式评估最终答案在不同平台和设备上的视觉呈现效果

- 持续审查公共基准测试和学术评估,判断其在 Perplexity 产品中的适用性,并将其适配并纳入我们的常规性能测量中

- 在一个小而高影响力的团队中工作,你的评估指标会直接影响产品变更,与技术领导紧密合作,衡量并提升答案质量

要求

- 技术领域的博士或硕士学历,或同等经验

- 4 年以上数据科学或机器学习经验

- 精通 Python 和 SQL(需编写生产级代码)

- 具有在现代云数据栈中工作的经验,特别是 AWS 和 Databricks

- 熟悉代理编码工作流,能够使用 AI 辅助开发工具加快迭代速度

优先考虑

- 1 年以上大规模使用 LLM 的经验,特别是 LLM 作为评判者的设置

- 具有面向客户网络产品或消费者应用的经验,具备大规模真实用户流量的背景

- 强大的研究背景,有将研究方法应用于实际 ML 问题的经验

- 有定义评估指标(如事实一致性、幻觉率、检索精度)和构建真实数据集的经验

查看英文原文

Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and our specialized data sources. We aim to use the latest models as they are released, but the intelligence frontier is a jagged one, and popular benchmarks do not effectively cover our use cases. In this role, you will build specialized evals to improve answer quality across Perplexity, covering search-based LLM answers and other scenarios popular with our users.

RESPONSIBILITIES

- Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness

- Design evaluation sets and methods specifically to measure the impact of tool calls (particularly web search retrieval) on the final answer's quality

- Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices

- Continuously review public benchmarks and academic evaluations for their applicability to the Perplexity product, adapting and incorporating them into our regular performance measurements

- Operate within a small, high-impact team where your evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve Answer Quality

QUALIFICATIONS

- PhD or MS in a technical field or equivalent experience

- 4+ years of experience in data science or machine learning

- Strong proficiency in Python and SQL (expected to write production-grade code)

- Experience building within a modern cloud data stack, specifically AWS and Databricks

- Comfortable with agentic coding workflows and using AI-assisted development tools to iterate faster

PREFERRED QUALIFICATIONS

- 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups

- Prior experience working on customer-facing web products or consumer apps, with real user traffic at scale

- A strong research background, with experience applying research methods to real-world ML problems

- Experience defining evaluation metrics (e.g., factual consistency, hallucination rate, retrieval precision) and building ground truth datasets

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位