远程工作雷达

技术成员(答案质量与评估)

Member of Technical Staff (Answer Quality & Evals)

AI未标注地域
公司Perplexity
薪资$200,000 - $350,000
工作地点San Francisco / Palo Alto
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-04-13
数据来源Ashby
前往企业招聘页投递 →

Perplexity 通过以大语言模型为核心的搜索引擎和专业数据源,每天为数千万用户提供可靠、高质量的答案。Answer Quality 团队确保我们的提示、工具、搜索系统、数据集和模型协同工作,为用户提供最佳体验。

随着我们的产品和代理功能不断发展,我们需要快速、可靠、贴近生产环境且可操作的评估系统。在这个职位中,你将构建和改进支持 Perplexity 答案质量的技术基础。这包括我们共享的评估基础设施以及用于重放和分析代理追踪的平台。你将与数据科学家、工程师和产品团队紧密合作,识别质量问题,衡量其影响,并将评估结果转化为产品改进。

职责

- 构建共享的评估基础设施,帮助团队运行可靠的评估、分析结果并做出产品和模型决策

- 开发用于重放和分析代理追踪的平台,以重现生产行为并诊断故障

- 构建并运营可扩展的系统,用于处理、存储和监控交互、追踪和评估数据

- 与数据科学家、工程师和产品团队合作,将答案质量问题转化为评估、分析和产品改进

- 在一个小而高影响力的团队中工作,你的工作将直接影响 Perplexity 如何衡量和提升答案质量

要求

- 4 年以上软件、数据或机器学习工程经验,有交付和运维生产系统的经历

- 精通 Python 和 SQL,具备系统设计、数据建模和分布式系统的基础知识

- 具备构建大数据系统的经验,包括分布式计算、大规模存储和高吞吐量流水线

- 在模糊的技术项目中展示出从初始设计到生产运维的主导能力

- 能够与数据科学家、工程师和产品合作伙伴高效协作

优先考虑

- 具备构建评估、实验、可观测性或机器学习基础设施的经验

- 熟悉大语言模型和代理系统,包括工具使用、执行追踪、重放和模拟

- 具备在大规模数据处理平台(如 Databricks、Snowflake 或 ClickHouse)上构建系统的经验

查看英文原文

Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and specialized data sources. The Answer Quality team ensures that our prompts, tools, search systems, datasets, and models work together to create the best possible experience for our users.

As our product and agent capabilities evolve, we need evaluation systems that are fast, reliable, production-faithful, and actionable. In this role, you will build and improve the technical foundations that support Answer Quality across Perplexity. This includes our shared evaluation infrastructure and the platform used to replay and analyze agent traces. You will work closely with data scientists, engineers, and product teams to identify quality problems, measure their impact, and turn evaluation findings into product improvements.

RESPONSIBILITIES

- Build shared evaluation infrastructure that helps teams run reliable evals, analyze results, and make product and model decisions

- Develop the platform for replaying and analyzing agent traces to reproduce production behavior and diagnose failures

- Build and operate scalable systems for processing, storing, and monitoring interaction, trace, and evaluation data

- Partner with data scientists, engineers, and product teams to turn answer-quality problems into evaluations, analyses, and product improvements

- Operate in a small, high-impact team where your work directly shapes how Perplexity measures and improves Answer Quality

QUALIFICATIONS

- 4+ years of software, data, or machine learning engineering experience shipping and operating production systems

- Strong proficiency in Python and SQL, with solid fundamentals in system design, data modeling, and distributed systems

- Experience building big-data systems, including distributed compute, large-scale storage, and high-volume pipelines

- Demonstrated ownership of ambiguous technical projects from initial design through production operation

- Ability to work effectively with data scientists, engineers, and product partners

PREFERRED QUALIFICATIONS

- Experience building evaluation, experimentation, observability, or machine learning infrastructure

- Familiarity with LLM and agent systems, including tool use, execution traces, replay, and simulation

- Experience building on top of large-scale data processing platforms such as Databricks, Snowflake, or ClickHouse

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位