远程工作雷达

用户研究员,AI评估

User Researcher, AI Evaluations

AI职能支持未标注地域
公司Notion
薪资未公开
工作地点San Francisco, California / New York, New York
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-06-18
数据来源Ashby
前往企业招聘页投递 →

我们是谁

Notion 是一个协作型 AI 工作空间,团队和代理可以一起思考 https://www.youtube.com/watch?v=vkpYpWfEK5s。我们正在打造一个地方,让您的知识、项目、会议和 AI 工具并存,使工作更快、更清晰、更不零散。数以百万计的个人、小型团队和大型公司都在 Notion 上开展工作。

Notinos(我们的员工)是实现这一未来工作的首批用户。我们重视工艺,打造持久的事物,并相信优秀的工作本质上仍是人类的。我们的目标不是推出下一个功能。每一位 Notinos 团队都在努力设定人类在 AI 时代协作的标准。从构建企业的记录系统到创建和管理 AI 代理,再到自动化处理琐碎的工作,我们非常关注为客户腾出更多时间去做他们人生中的重要工作。

职位简介:

我们正在寻找一位经验丰富的 UX 研究员,来定义和扩展我们评估 Notion AI 驱动体验的方法——重点不仅在于模型输出质量,还在于用户从发现、设定目标、委派任务、审查结果到与 AI 建立信任的端到端产品体验。

这个职位位于研究技能和评估运营的交汇点:你将进行研究,揭示用户的思维模型、期望和失败/恢复行为,然后将这些洞察转化为可重复使用的评估标准、工作流程和测量方法,供产品、设计、工程和数据科学团队一致应用。

该职位可在旧金山或纽约市任一地点办公。我们在周一、周二和周四(我们的锚点日)在办公室工作,因为我们在现场一起思考和构建时表现最佳。我们希望找到一位在这些日子能与团队一起工作的候选人。

你将取得的成果:

- 定义“良好”的标准(框架和评估标准):建立清晰、可重复使用的评估标准,反映真实用户期望——有用性、信任度、语气、控制感和透明度。你将把定性洞察转化为可跨团队和时间一致应用的评分指导。

- 运行定期评估(纵向和特定功能):运行定期的纵向和特定功能的调查和研究,根据定义的评估标准衡量体验质量。领导定性研究、并排比较和人工介入评估

查看英文原文

WHO WE ARE

Notion is the collaborative AI workspace where teams and agents think together https://www.youtube.com/watch?v=vkpYpWfEK5s. We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion.

Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work.

ABOUT THE ROLE:

We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for the end-to-end product experience where people discover, set goals, delegate work, review results, and build trust over time with AI.

This role sits at the intersection of research craft and evaluation operations: you’ll run studies that uncover user mental models, expectations, and failure/recovery behaviors, then translate those insights into reusable rubrics, workflows, and measurement approaches that product, design, engineering, and data science can apply consistently.

This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days.

WHAT YOU'LL ACHIEVE:

- Define what “good” looks like (frameworks & rubrics): Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness, trust, tone, control, and transparency. You’ll translate qualitative insight into scoring guidance that can be applied consistently across teams and over time.

- Run recurring evals (longitudinal & feature-specific): Run recurring longitudinal and feature-specific surveys and studies to measure experience quality over time against defined rubrics. Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluation efforts to deepen understanding of where experiences break down and how they can improve. You’ll help teams spot regressions, benchmark improvements, and understand when expectations shift.

- Anchor evaluation in real workflows (context > isolated feedback): Ensure evals reflect jobs-to-be-done, user intent, and the full interaction journey (goal setting, delegation, review, iteration), not just decontextualized thumbs up/down. You’ll help teams understand who is evaluating, what they’re trying to do, and why outputs succeed or fail.

- Identify failure modes & recovery behavior (guardrails): Uncover breakdowns, regressions, and edge cases across the system—from model behavior to UI and integrations—and study how people notice issues, correct them, and continue their work. You’ll turn these insights into actionable guidance for guardrails, fixes, and prioritization.

- Operationalize evaluation with partners (process & tooling): Collaborate closely with Product, Design, Engineering, and Data Science to align on target use cases and build scalable evaluation loops (human-in-the-loop review, longitudinal studies, and calibration of automated/LLM-judge approaches against human judgment).

SKILLS YOU'LL NEED TO BRING:

- Ability to operationalize insight into measurement: You’re comfortable turning “soft” user expectations (trust, tone, usefulness, clarity) into concrete rubrics, scoring guidelines, and observable metrics.

- AI fluency and systems thinking: You’re curious and hands-on with AI products, and can reason about how model behavior, uncertainty, and system constraints shape user experience. You also have experience evaluating AI-enabled products (LLMs, agents, generative UI/workflow automation) and working with Data Science/ML partners on measurement strategy and evaluation tooling.

- Clear communication and impact orientation: You can align diverse partners around shared definitions of quality and create artifacts that enable teams to act consistently. You tailor storytelling to different audiences, connect research to business outcomes, and drive follow-through so insights translate into product change.

- Strong UX research craft (quant + qual): You can choose the right methods for the question— interviews, benchmarking, surveys, experiments—and synthesize into actionable guidance. You also can prioritize ruthlessly, work through ambiguity, and balance scrappy iteration with deep dives when needed.

- Pragmatism in fast-moving environments: You can prioritize ruthlessly, work through ambiguity, and balance scrappy iteration with deep dives when needed.

- Experience: 5+ years doing UX research in industry

NICE TO HAVES:

- Familiarity with LLM-as-judge methods, prompt design for evaluators, or “golden dataset” creation

- Experience using AI research tooling for rapid synthesis and communication (e.g., Dovetail, Listen Labs, Maze, Outset, etc.), as well as AI observability tooling like Braintrust

- Experience using data querying languages (e.g., SQL), scripting languages (e.g., Python), or statistical/mathematical software (e.g., R, SAS, Matlab, etc.)

- Master’s or PhD in HCI, Psychology, Behavioral Science, Anthropology, Sociology, or a related field

- You’re familiar with the work of computing heroes like Douglas Engelbart, Alan Kay, Bret Victor, etc. — and understand why we're big fans.

Notion is committed to providing highly competitive cash compensation, equity, and benefits. The compensation offered for this role will be based on multiple factors such as location, the role’s scope and complexity, and the candidate’s experience and expertise, and may vary from the range provided below. For roles based in San Francisco or New York City, the estimated base salary range for this role is $196,000-$230,000 per year.

By clicking “Submit Application”, I understand and agree that Notion and its affiliates and subsidiaries will collect and process my information in accordance with Notion’s Global Recruiting Privacy Policy https://notion.notion.site/Notion-Global-Recruiting-Privacy-Policy-fc3eb4e829354a26a2bb6fd5e289b550?pvs=74 and NYLL 144 https://notion.notion.site/Ashby-AI-Bias-Audit-2b0efdeead05803bbbfae159ec86c528.

#LI-Onsite

A NOTE ON AI

You don’t need deep AI expertise for every role, but we do expect every Notino to be intellectually curious, drawn to tinkering and discovery, and excited to use AI as a real collaborator in their work. For some roles, AI fluency is a core requirement — when that’s the case, we'll say so explicitly in the qualifications. People who thrive here don’t treat AI as a novelty. They use it to think better, and make their work easier for others to build on.

EQUAL OPPORTUNITY & ACCOMMODATIONS

We hire talented people from a wide range of backgrounds. If you’re excited about this role but don’t meet every bullet, we still encourage you to apply. Notion is an equal opportunity employer and does not discriminate on the basis of any legally protected characteristic. Consistent with applicable law, we will consider for employment qualified applicants with arrest and conviction records. Notion provides reasonable accommodations during the application process; if you need one, please let your recruiter know.

Notion is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristic. Notion considers qualified applicants with criminal histories, consistent with applicable federal, state and local law. Notion is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please let your recruiter know.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

现场营销人员

NotionSan Francisco, California / New York, New FullTime11 天前
市场运营未标注地域

← 返回全部职位