AI 安全专员 - 远程 | 最高 84 美元/小时
AI Safety Specialist - Remote | Upto $84/hr
Mercor将顶尖的创意和技术人才与领先的AI研究实验室联系起来。公司总部位于旧金山,我们的投资方包括Benchmark、General Catalyst、Peter Thiel、Adam D'Angelo、Larry Summers和Jack Dorsey。
职位:AI安全红队成员
类型:合同工
薪酬:每小时70-84美元
地点:远程工作
岗位职责
- 设计对抗性提示,对前沿AI模型进行压力测试。
- 识别越狱行为、不安全行为、幻觉和政策失效。
- 在虚假信息、网络、生物安全、欺诈、政治内容和其他敏感领域评估模型的鲁棒性。
- 记录漏洞,并为安全基准测试和红队报告做出贡献。
- 与AI研究人员合作,提升模型对齐度、鲁棒性和安全性。
资格要求
必须具备
- 计算机科学、网络安全、新闻学、传播学、心理学、生物学、化学、公共政策或相关专业的学士学位或更高学历。
- 在AI安全、AI红队、信任与安全、网络安全、调查新闻、生命科学或相关领域有5年以上专业经验。
- 具备强大的分析推理、提示设计和书面沟通能力。
- 有设计对抗性提示或评估前沿AI系统的经验。
优先考虑
- 有AI红队、RLHF、SFT、AI对齐或信任与安全方面的经验。
- 熟悉越狱测试、提示工程或对抗性评估方法。
- 在网络、生物安全、政治内容、虚假信息或科学安全等一个或多个灰色领域具有专业知识。
申请流程(需要20-30分钟完成)
- 上传简历
- 基于简历的AI面试
- 提交表单
资源与支持
- 有关面试流程和平台信息的详细信息,请查看:
- 如有任何帮助或支持需求,请联系:
备注:我们的团队每天都会审核申请。请完成AI面试和申请步骤,以考虑此机会。
最初发布于Himalayas
查看英文原文
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.
Position: AI Safety Red Teamer
Type:Contract
Compensation:$70–$84/hour
Location:Remote
Role Responsibilities
- Design adversarial prompts to stress-test frontier AI models.
- Identify jailbreaks, unsafe behaviors, hallucinations, and policy failures.
- Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains.
- Document vulnerabilities and contribute to safety benchmarking and red-teaming reports.
- Collaborate with AI researchers to improve model alignment, robustness, and safety.
Qualifications
Must-Have
- Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline.
- 5+ years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related field.
- Strong analytical reasoning, prompt design, and written communication skills.
- Experience designing adversarial prompts or evaluating frontier AI systems.
Preferred
- Experience with AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety.
- Familiarity with jailbreak testing, prompt engineering, or adversarial evaluation methodologies.
- Expertise in one or more grey-area domains, including cyber, biosecurity, political content, misinformation, or scientific safety.
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to:
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
Originally posted on Himalayas