AI 评估员:评估购物助手
AI Evaluators: Assessing A Shopping Assistant
WHAT WE'RE RESEARCHING
我们正在招聘AI评估人员,以评估新型数字购物助手的准确性和帮助性。该项目旨在了解系统如何处理现实中的电子商务查询,以及其逻辑在哪些方面存在不足。您的分析将直接用于改进底层模型和响应质量。
HOW IT WORKS
您将在我们的定制平台上查看用户与购物助手之间的实际交互记录。在分析这些对话时,您需要识别具体的故障、逻辑错误或无帮助的产品推荐。随后,您将创建结构化的评分标准和验证器,以一致地评估未来的响应质量。这是一项持续的远程工作,每周需投入20小时以上。
WHO THIS IS FOR
此机会适合质量保障专家、AI数据评估员和电子商务专业人士,要求具备细致的观察力。我们欢迎有提示工程、复杂数据标注或软件测试经验的申请者。您应能够深入分析文本交互,并从零开始构建结构化的评估框架。
WHAT YOU'LL DO
- 审查AI购物助手的真实用户交互记录
- 识别文本中的逻辑故障、不准确之处或不良推荐
- 创建结构化的评分标准和验证器以评估响应质量
- 承诺每周在我们的内部平台上进行20小时以上的评估工作
WHO SHOULD APPLY
- 具有数据评估、质量保障或AI训练的经验
- 强大的分析能力,能够发现文本中的细微错误
- 熟悉电子商务搜索和数字购物体验
- 能够承诺每周持续投入20小时以上的工作量
COMPENSATION
每小时50美元
READY TO PARTICIPATE?
现在开始您的付费面试 https://terac.com/interview/start/r/c5aab5ae-fa85-45ab-918d-a8b7cdf9f0ec?utm_source=ashby_listing_description
ABOUT TERAC
Terac 正在打造全球最大的经过认证的人类专家池,用于AI。研究人员、AI实验室和产品团队使用 Terac 在各个行业、语言和技能领域招募、筛选和支付研究参与者。
了解更多请访问 terac.com https://terac.com 或在 YouTube 上关注 @jointerac https://www.youtube.com/@jointerac.
查看英文原文
WHAT WE'RE RESEARCHING
We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.
HOW IT WORKS
You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.
WHO THIS IS FOR
This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.
WHAT YOU'LL DO
- Review real user interaction traces with an AI shopping assistant
- Identify logical failures, inaccuracies, or poor recommendations in the text
- Create structured rubrics and verifiers to judge response quality
- Commit to 20+ hours per week of evaluation work on our internal platform
WHO SHOULD APPLY
- Experience in data evaluation, quality assurance, or AI training
- Strong analytical skills with the ability to spot subtle errors in text
- Familiarity with e-commerce search and digital shopping experiences
- Ability to commit to a sustained workload of 20+ hours per week
COMPENSATION
$50 per hour
READY TO PARTICIPATE?
Start your paid interview now https://terac.com/interview/start/r/c5aab5ae-fa85-45ab-918d-a8b7cdf9f0ec?utm_source=ashby_listing_description
ABOUT TERAC
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at terac.com https://terac.com or on YouTube at @jointerac https://www.youtube.com/@jointerac.