远程工作雷达

AI安全专家 - 对抗机器学习

AI Safety Expert - Adversarial ML

AI开发工程限定地区(需当地身份)
公司mercor
薪资$16 - $22
工作地点Singapore
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Contractor
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 Singapore 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

关于职位
Mercor 将顶尖的创意和技术人才与领先的 AI 研究实验室联系起来。公司总部位于旧金山,我们的投资人包括 Benchmark、General Catalyst、Peter Thiel、Adam D'Angelo、Larry Summers 和 Jack Dorsey。
职位:AI 安全专家 — 英语 & 泰米尔语
类型:合同制
薪资:16–22 美元/小时
地点:远程

岗位职责

  • 通过 jailbreak、prompt injection 和滥用案例对对话 AI 模型和代理进行红队测试。识别偏见利用和多轮操纵。
  • 通过标注失败、分类漏洞和标记系统性风险,生成高质量的人类数据。
  • 通过遵循分类法、基准测试和操作手册来建立结构,确保测试的一致性。
  • 通过生成客户可采取行动的报告、数据集和攻击案例,实现可重复记录。
  • 独立且异步工作,以满足截止日期并提升 AI 模型性能。

任职要求

必须具备

  • 流利的语言技能:英语 & 泰米尔语。需要英语和泰米尔语的母语级流利。
  • 对语言和内容的准确性、完整性和适当性有良好的判断力。
  • 对细微错误、不一致和漏洞有严谨的关注。
  • 结构化的工作方法,遵守指南和质量标准。
  • 与技术人员和非技术人员清晰沟通。
  • 在项目、任务类型和客户之间具有适应能力。

优先考虑

  • 攻击性机器学习经验:jailbreak 数据集、prompt injection、RLHF/DPO 攻击、模型提取。
  • 网络安全技能:渗透测试、漏洞开发、逆向工程。
  • 社会技术风险知识:骚扰/虚假信息探测、滥用分析、对话 AI 测试。
  • 创造性探测技能:心理学、表演、为非常规对抗思维写作。

申请流程(需 20–30 分钟完成)

  • 上传简历
  • 基于简历的 AI 面试
  • 提交表单

资源与支持

  • 有关面试流程和平台信息的详细信息,请查看:
  • 如有任何帮助或支持,请联系:

备注:我们的团队每天都会审核申请。请完成 AI 面试和申请步骤,以考虑此机会。
最初发布于 Himalayas

查看英文原文

About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.
Position: AI Safety Experts — English & Tamil
Type:Contract
Compensation:$16–$22/hour
Location:Remote
Role Responsibilities

  • Red team conversational AI models and agents through jailbreaks, prompt injections, and misuse cases. Identify bias exploitation and multi-turn manipulation.
  • Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing.
  • Document reproducibly by producing reports, datasets, and attack cases that customers can act on.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.

Qualifications

Must-Have

  • Fluent Language Skills Required: English & Tamil. Native fluency in English and Tamil is required.
  • Strong judgment about language and content accuracy, completeness, and appropriateness.
  • Rigorous attention to subtle errors, inconsistencies, and gaps.
  • Structured approach to work, adhering to guidelines and quality standards.
  • Clear communication with both technical and non-technical audiences.
  • Adaptability across projects, task types, and customers.

Preferred

  • Experience in Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
  • Cybersecurity skills: penetration testing, exploit development, reverse engineering.
  • Knowledge of socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.
  • Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

  • For details about the interview process and platform information, please check:
  • For any help or support, reach out to:

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位