技术成员(数据科学家,评估)
Member of Technical Staff (Data Scientist, Evals)
Perplexity 通过以 LLM 为主导的搜索引擎和我们专业的数据源,每天为数千万用户提供可靠、高质量的答案。我们希望在最新模型发布时立即使用,但智能边界是不规则的,流行的基准测试并不能有效覆盖我们的使用场景。在这个职位中,你将构建专门的评估方案,以提升 Perplexity 的答案质量,涵盖基于搜索的 LLM 答案和其他用户常用场景。
职责
- 架构并维护自动化评估流程,以评估 Perplexity 产品中的答案质量,确保准确性和帮助性达到高标准
- 设计专门的评估集和方法,以衡量工具调用(尤其是网络搜索检索)对最终答案质量的影响
- 开发基于视觉语言模型(VLM)的解决方案,以程序化方式评估最终答案在不同平台和设备上的视觉呈现效果
- 持续审查公共基准测试和学术评估,判断其在 Perplexity 产品中的适用性,并将其适配并纳入我们的常规性能测量中
- 在一个小而高影响力的团队中工作,你的评估指标会直接影响产品变更,与技术领导紧密合作,衡量并提升答案质量
要求
- 技术领域的博士或硕士学历,或同等经验
- 4 年以上数据科学或机器学习经验
- 精通 Python 和 SQL(需编写生产级代码)
- 具有在现代云数据栈中工作的经验,特别是 AWS 和 Databricks
- 熟悉代理编码工作流,能够使用 AI 辅助开发工具加快迭代速度
优先考虑
- 1 年以上大规模使用 LLM 的经验,特别是 LLM 作为评判者的设置
- 具有面向客户网络产品或消费者应用的经验,具备大规模真实用户流量的背景
- 强大的研究背景,有将研究方法应用于实际 ML 问题的经验
- 有定义评估指标(如事实一致性、幻觉率、检索精度)和构建真实数据集的经验
查看英文原文
Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and our specialized data sources. We aim to use the latest models as they are released, but the intelligence frontier is a jagged one, and popular benchmarks do not effectively cover our use cases. In this role, you will build specialized evals to improve answer quality across Perplexity, covering search-based LLM answers and other scenarios popular with our users.
RESPONSIBILITIES
- Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness
- Design evaluation sets and methods specifically to measure the impact of tool calls (particularly web search retrieval) on the final answer's quality
- Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices
- Continuously review public benchmarks and academic evaluations for their applicability to the Perplexity product, adapting and incorporating them into our regular performance measurements
- Operate within a small, high-impact team where your evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve Answer Quality
QUALIFICATIONS
- PhD or MS in a technical field or equivalent experience
- 4+ years of experience in data science or machine learning
- Strong proficiency in Python and SQL (expected to write production-grade code)
- Experience building within a modern cloud data stack, specifically AWS and Databricks
- Comfortable with agentic coding workflows and using AI-assisted development tools to iterate faster
PREFERRED QUALIFICATIONS
- 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups
- Prior experience working on customer-facing web products or consumer apps, with real user traffic at scale
- A strong research background, with experience applying research methods to real-world ML problems
- Experience defining evaluation metrics (e.g., factual consistency, hallucination rate, retrieval precision) and building ground truth datasets