远程工作雷达

软件工程师:AI评估任务付费面试

Software Engineers: Paid Interview on AI Evaluation Tasks

AI开发工程未标注地域
公司Terac
薪资未公开
工作地点United States
地域资格未标注地域
时区要求无特别要求
用工类型Contract
发布时间今天
数据来源Ashby
前往企业招聘页投递 →

我们正在研究

我们正在进行一项付费研究,评估用于测试AI代理的编码环境的质量和真实性。我们正在开发一套全面的编程任务和评估工具包,以准确衡量AI性能。您的反馈将直接影响这些代理与现实世界工程标准的基准测试方式。

如何参与

在此次远程会议中,您将审查多个编程任务及其对应的评估工具包。您将评估每个编码挑战的技术准确性、复杂性和真实性。我们会要求您讲解评估环境的逻辑并识别任何潜在缺陷。最后,您将提供这些环境与行业标准实践相比的反馈意见。

适合人群

我们寻找有实际经验构建、审查或测试真实编程任务的软件工程师。我们欢迎熟悉评估工具包的全栈开发人员、后端工程师、测试自动化工程师和系统架构师。理想的候选人了解使编码挑战具有鲁棒性、可验证性和技术性的要素。

您将负责

- 审查和评估真实编程任务的质量
- 评估用于AI测试的编码环境和技术工具包
- 向我们讲解提供的代码结构中的潜在缺陷
- 提供对挑战难度和真实性的反馈意见

申请条件

- 具有软件工程师或开发人员的专业经验
- 熟悉构建或验证编程任务
- 具有代码审查和评估工具包的经验
- 能够讨论技术架构和测试方法

报酬

每小时75美元

是否准备好参与?

立即开始您的付费面试 https://terac.com/interview/start/r/900229bf-5f7d-49c2-b927-1054bd745f0e?utm_source=ashby_listing_description

关于TERAC

Terac正在打造全球最大的经过认证的人类专家池,用于AI。研究人员、AI实验室和产品团队使用Terac在各个行业、语言和技能水平中招募、筛选和支付研究参与者。

了解更多请访问 terac.com https://terac.com 或在YouTube上关注 @jointerac https://www.youtube.com/@jointerac

查看英文原文

WHAT WE'RE RESEARCHING

We're running a paid study on the quality and realism of coding environments designed to test AI agents. We are developing a comprehensive suite of programming tasks and evaluation harnesses to measure AI performance accurately. Your feedback directly shapes how these agents are benchmarked against real-world engineering standards.

HOW IT WORKS

During this remote session, you will review several programming tasks and their corresponding evaluation harnesses. You will assess the technical accuracy, complexity, and realism of each coding challenge. We will ask you to walk through the logic of the evaluation environments and identify any potential flaws. Finally, you will provide feedback on how these environments compare to standard industry practices.

WHO THIS IS FOR

We are looking for software engineers who have hands-on experience building, reviewing, or testing realistic programming tasks. We welcome full-stack developers, backend engineers, test automation engineers, and systems architects who are familiar with evaluation harnesses. Ideal candidates understand what makes a coding challenge robust, verifiable, and technically sound.

WHAT YOU'LL DO

- Review and assess the quality of realistic programming tasks

- Evaluate coding environments and technical harnesses for AI testing

- Walk us through potential flaws in the provided code structures

- Provide feedback on the difficulty and realism of the challenges

WHO SHOULD APPLY

- Professional experience as a software engineer or developer

- Familiarity with building or verifying programming tasks

- Experience with code review and evaluation harnesses

- Comfortable discussing technical architecture and testing methodologies

COMPENSATION

$75 per hour

READY TO PARTICIPATE?

Start your paid interview now https://terac.com/interview/start/r/900229bf-5f7d-49c2-b927-1054bd745f0e?utm_source=ashby_listing_description

ABOUT TERAC

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

Learn more at terac.com https://terac.com or on YouTube at @jointerac https://www.youtube.com/@jointerac.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位