AI安全专员 - 双语
AI Safety Specialist - Bilingual
职位简介
Mercor将顶尖的创意和技术人才与领先的AI研究实验室联系起来。公司总部位于旧金山,我们的投资者包括Benchmark、General Catalyst、Peter Thiel、Adam D'Angelo、Larry Summers和Jack Dorsey。
职位:AI安全专家 — 英语和马拉雅拉姆语
类型:合同工
薪酬:16–22美元/小时
地点:远程工作
岗位职责
- 对对话式AI模型和代理进行红队测试,以识别越狱、提示注入和滥用案例。
- 通过标注失败案例、分类漏洞和标记系统性风险,生成高质量的人类数据。
- 按照分类法、基准测试和操作手册进行结构化工作,确保测试的一致性。
- 可重复记录,生成客户可采取行动的报告、数据集和攻击案例。
- 独立且异步工作,按时完成任务,同时提升AI模型性能。
- 任职要求
必须具备
- 流利的语言技能:英语和马拉雅拉姆语。需要英语和马拉雅拉姆语的母语水平。
- 对语言和内容准确性有良好的判断力。
- 对细节有严谨的关注,并能发现细微错误。
- 按照指南和质量标准进行结构化的工作方法。
- 与技术人员和非技术人员沟通清晰的能力。
- 在不同项目、任务类型和客户之间具备适应能力。
- 优先考虑
- 有对抗性机器学习经验:越狱数据集、提示注入、RLHF/DPO攻击、模型提取。
- 网络安全技能:渗透测试、漏洞开发、逆向工程。
- 对社会技术风险的理解:骚扰/虚假信息探测、滥用分析、对话式AI测试。
- 创造性探测技能:心理学、表演、为非常规对抗思维写作。
- 申请流程(需20–30分钟完成)
- 上传简历
- 基于简历的AI面试
- 提交表单
- 资源与支持
- 有关面试流程和平台信息的详细信息,请查看:
- 如有任何帮助或支持需求,请联系:
备注:我们团队每天都会审核申请。请完成AI面试和申请步骤,以便考虑此机会。
最初发布于Himalayas
查看英文原文
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.
Position: AI Safety Experts — English & Malayalam
Type:Contract
Compensation:$16–$22/hour
Location:Remote
Role Responsibilities
- Red team conversational AI models and agents to identify jailbreaks, prompt injections, and misuse cases.
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
- Apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing.
- Document reproducibly by producing reports, datasets, and attack cases that customers can act on.
- Work independently and asynchronously to meet deadlines while improving AI model performance.
Qualifications
Must-Have
- Fluent Language Skills Required:English & Malayalam. Native fluency in English and Malayalam is required.
- Strong judgment about language and content accuracy.
- Rigorous attention to detail and ability to notice subtle errors.
- Structured approach to work following guidelines and quality standards.
- Clear communication skills for technical and non-technical audiences.
- Adaptability across projects, task types, and customers.
Preferred
- Experience in Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
- Cybersecurity skills: penetration testing, exploit development, reverse engineering.
- Understanding of socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.
- Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to:
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
Originally posted on Himalayas