Anthropic研究员项目,人工智能安全与安全
Anthropic Fellows Program, AI Safety & Security
Anthropic 介绍
Anthropic 的使命是创造可靠、可解释且可引导的 AI 系统。我们希望 AI 对我们的用户以及整个社会都是安全且有益的。我们的团队是一支快速发展的由致力于研究、工程、政策专家和商业领袖组成的团队,共同构建有益的 AI 系统。
目前申请已关闭,我们将在此页面重新开放时进行更新。
Anthropic Fellows 计划概述
Anthropic Fellows 计划旨在培养 AI 研究和工程人才。我们为有潜力的技术人才提供资金和指导,无论其以往经验如何。
Fellows 将主要使用外部基础设施(例如开源模型、公共 API)来开展与我们研究重点一致的实证项目,并以产出公开成果(例如论文提交)为目标。在我们早期的一个小组中,超过 80% 的 Fellow 产生了论文。我们每年运行多个 Fellow 小组,并按滚动方式审核申请。
你将获得什么
- 4 个月的全职研究时间
- 与 Anthropic 研究人员直接指导
- 可以使用共享办公空间(在加利福尼亚州伯克利或英国伦敦)
- 连接更广泛的 AI 安全与安全研究社区
- 每周津贴 3,850 美元 / 2,310 英镑 / 4,300 加元 + 福利(具体因国家而异)
- 计算资源资助(约 15,000 美元/月)和其他研究支出
面试流程
面试流程包括初步申请及推荐人核查、技术评估及面试,以及研究讨论。
我们鼓励你申请,即使你认为自己并不符合每一个条件。并非所有优秀的候选人会满足所有列出的条件。研究表明,来自代表性不足群体的人更容易产生“冒名顶替综合症”并质疑自己的竞争力,因此我们敦促你不要过早排除自己,如果你对这项工作感兴趣,请提交申请。我们认为我们正在构建的 AI 系统具有巨大的社会和伦理影响。我们认为这使代表性更加重要,我们努力在团队中包含多样化的观点。
薪酬
该职位的预期基本津贴为每周 3,850 美元 / 2,310 英镑 / 4,300 加元,预计每周工作 40 小时,持续 4 个月(可能延长)。
Fellows 工作流
由于成功的原因
查看英文原文
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
Applications are currently closed, we will update this page when they re-open.
Anthropic Fellows Program overview
The Anthropic Fellows Program is designed to foster AI research and engineering talent. We provide funding and mentorship to promising technical talent - regardless of previous experience.
Fellows will primarily use external infrastructure (e.g. open-source models, public APIs) to work on an empirical project aligned with our research priorities, with the goal of producing a public output (e.g. a paper submission). In one of our earlier cohorts, over 80% of fellows produced papers. We run multiple cohorts of Fellows each year and review applications on a rolling basis.
What to expect
- 4 months of full-time research
- Direct mentorship from Anthropic researchers
- Access to a shared workspace (in either Berkeley, California or London, UK)
- Connection to the broader AI safety and security research community
- Weekly stipend of 3,850 USD / 2,310 GBP / 4,300 CAD + benefits (these vary by country)
- Funding for compute (~$15k/month) and other research expenses
Interview process
The interview process will include an initial application & reference check, technical assessments & interviews, and a research discussion.
We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.
Compensation
The expected base stipend for this role is 3,850 USD / 2,310 GBP / 4,300 CAD per week, with an expectation of 40 hours per week for 4 months (with possible extension).
Fellows workstreams
Due to the success of the Anthropic Fellows for AI Safety Research program, we are now expanding it across teams at Anthropic. We expect there to be significant overlap in the types of skills and responsibilities across the roles and will by default consider candidates for all the workstreams.
Some of the workstreams may include unique assessment steps; we therefore ask you for workstream preferences in the application. You can see an overview of the current workstreams below:
- AI Safety Fellows
- AI Security Fellows
- ML Systems & Performance Fellows
- Reinforcement Learning Fellows
- The Anthropic Institute Fellows (Economics & Policy)
This page is specific to one of the Anthropic Fellows Workstreams, see also the main Anthropic Fellows posting.
Across the workstreams, you may be a good fit if you:
- Are motivated by making sure AI is safe and beneficial for society as a whole
- Are excited to transition into empirical AI research and would be interested in a full-time role at Anthropic
- Have a strong technical background in computer science, mathematics, or physics
- Thrive in fast-paced, collaborative environments
- Can implement ideas quickly and communicate clearly
Strong candidates may also have:
- Strong background in a discipline relevant to a specific Fellows workstream (e.g. economics, social sciences, or cybersecurity)
- Experience in areas of research or engineering related to their workstream
Candidates must be:
- Fluent in Python programming
- Available to work full-time on the Fellows program
AI Safety Fellows
Mentors, research areas, & past projects
Fellows will undergo a project selection & mentor matching process. Potential mentors include:
- Sam Bowman
- Sara Price
- Alex Tamkin
- Nina Panickssery
- Trenton Bricken
- Logan Graham
- Jascha Sohl-Dickstein
- Joe Benton
- Collin Burns
- Fabien Roger
- Samuel Marks
- Kyle Fish
- Ethan Perez
Our mentors will lead projects in select AI safety research areas, such as:
- Scalable Oversight: Developing techniques to keep highly capable models helpful and honest, even as they surpass human-level intelligence in various domains.
- Adversarial Robustness and AI Control: Creating methods to ensure advanced AI systems remain safe and harmless in unfamiliar or adversarial scenarios.
- Model Organisms: Creating model organisms of misalignment to improve our empirical understanding of how alignment failures might arise.
- Model Internals / Mechanistic Interpretability: Advancing our understanding of the internal workings of large language models to enable more targeted interventions and safety measures.
- AI Welfare: Improving our understanding of potential AI welfare and developing related evaluations and mitigations.
On our Alignment Science and Frontier Red Team blogs, you can read about past projects, including:
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data: Alex Cloud and Minh Le, et al., mentors including Samuel Marks and Owain Evans
- Open-source circuits: Michael Hanna and Mateusz Piotrowski with mentorship from Emmanuel Ameisen and Jack Lindsey
For a full list of representative projects for each area, please see these blog posts: Introducing the Anthropic Fellows Program for AI Safety Research, Recommendations for Technical AI Safety Research Directions.
Unique candidate criteria
You might be a particularly great fit for this workstream if you:
- Are motivated by reducing catastrophic risks from advanced AI systems
- Have experience with empirical ML research projects
- Have experience working with large language models
- Have experience in one of the research areas mentioned above
- Have a track record of open-source contributions
AI Security Fellows
Mentors, research areas, & past projects
Fellows will undergo a project selection & mentor matching process. Potential mentors include:
- Nicholas Carlini
- Keri Warr
- Evyatar Ben Asher
- Keane Lucas
- Newton Cheng
On our Alignment Science and Frontier Red Team blogs, you can read about some past Fellows projects, including:
- AI agents find $4.6M in blockchain smart contract exploits: Winnie Xiao and Cole Killian, mentored by Nicholas Carlini and Alwin Peng
- Strengthening Red Teams: A Modular Scaffold for Control Evaluations: Chloe Loughridge et al., mentored by Jon Kutasov and Joe Benton
Unique candidate criteria
You might be a particularly great fit for this workstream if you:
- Are motivated by reducing catastrophic risks from advanced AI systems
- Have contributed to open-source projects in LLM- or security-adjacent repositories
- Have demonstrated success in bringing clarity and ownership to ambiguous technical problems
- Have experience with pentesting, vulnerability research, or other offensive security work
- Have a demonstrated willingness to do the "dirty work" that produces high-quality outputs
- Have reported CVEs or been awarded bug bounties
- Have experience with empirical ML research projects
- Have experience with deep learning frameworks and experiment management
Logistics
Logistics Requirements: To participate in the Fellows program, you must have work authorization in the US, UK, or Canada and be located in that country during the program.
Workspace Locations: We have designated shared workspaces in London and Berkeley where fellows will work from and mentors will visit. We are also open to remote fellows in the UK, US, or Canada. We will ask you about your availability to work from Berkeley or London (full- or part-time) during the program.
Visa Sponsorship: We are not currently able to sponsor visas for fellows. To participate in the Fellows program, you need to have or independently obtain full-time work authorization in the UK, the US, or Canada.
Program Duration: The program runs for 4 months, full-time. If you can't commit to the full duration, please still apply and note your constraints in the application. We review these requests on a case-by-case basis.
Please note: We do not guarantee that we will make any full-time offers to fellows. However, strong performance during the program may indicate that a Fellow would be a good fit for full-time roles at Anthropic. In previous cohorts, 25-50% of fellows received a full-time offer, and we’ve supported many more to go on to do great work on AI safety and security at other organizations.
Applications and interviews are managed by Constellation, our recruiting partner. Constellation also runs the Berkeley workspace and provides program support for fellows working on AI safety and security; fellows on capabilities-focused projects are supported directly by Anthropic. All applicants currently use the same application portal but we are working to separate applications for safety/security and capabilities-focused projects in future rounds.
Applications are currently closed, we will update this page when they re-open.
The below are Anthropic's policies for full-time roles. These do NOT apply to the Fellows Program.
Logistics
Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.
Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.
Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings.
How we're different
We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.
The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
Come work with us!
Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.