Summary
Anthropic is a public benefit corporation focused on creating reliable, interpretable, and steerable AI systems that are safe and beneficial for society. The Anthropic Fellows Program offers four months of full-time research and engineering work with mentorship, external infrastructure, and funding to pursue empirical projects in AI safety and security. Fellows may work on areas including scalable oversight, adversarial robustness, model internals, AI welfare, vulnerability research, and security.
Responsibilities
- 4 months of full-time research
- Direct mentorship from Anthropic researchers
- Access to a shared workspace (in either Berkeley, California or London, UK)
- Connection to the broader AI safety and security research community
- Funding for compute (~$15k/month) and other research expenses
- Developing techniques to keep highly capable models helpful and honest, even as they surpass human-level intelligence in various domains
- Creating methods to ensure advanced AI systems remain safe and harmless in unfamiliar or adversarial scenarios
- Creating model organisms of misalignment to improve our empirical understanding of how alignment failures might arise
- Advancing our understanding of the internal workings of large language models to enable more targeted interventions and safety measures
- Improving our understanding of potential AI welfare and developing related evaluations and mitigations
- Have contributed to open-source projects in LLM- or security-adjacent repositories
- Have demonstrated success in bringing clarity and ownership to ambiguous technical problems
- Have experience with pentesting, vulnerability research, or other offensive security work
- Have reported CVEs or been awarded bug bounties
- Have experience with empirical ML research projects
- Have experience with deep learning frameworks and experiment management
Skills
- * Are motivated by making sure AI is safe and beneficial for society as a whole
- * Are excited to transition into empirical AI research and would be interested in a full-time role at Anthropic
- * Have a strong technical background in computer science, mathematics, or physics
- * Thrive in fast-paced, collaborative environments
- * Can implement ideas quickly and communicate clearly
- * Fluent in Python programming
- * Available to work full-time on the Fellows program
- * Are motivated by reducing catastrophic risks from advanced AI systems
- * Have experience with empirical ML research projects
- * Have experience working with large language models
- * Have experience in one of the research areas mentioned above
- * Have a track record of open-source contributions
- * Are motivated by reducing catastrophic risks from advanced AI systems
- * Have contributed to open-source projects in LLM- or security-adjacent repositories
- * Have demonstrated success in bringing clarity and ownership to ambiguous technical problems
- * Have experience with pentesting, vulnerability research, or other offensive security work
- * Have a demonstrated willingness to do the "dirty work" that produces high-quality outputs
- * Have reported CVEs or been awarded bug bounties
- * Have experience with empirical ML research projects
- * Have experience with deep learning frameworks and experiment management
- To participate in the Fellows program, you must have work authorization in the US, UK, or Canada and be located in that country during the program
- We are also open to remote fellows in the UK, US, or Canada
- To participate in the Fellows program, you need to have or independently obtain full-time work authorization in the UK, the US, or Canada
- The program runs for 4 months, full-time
- * Strong background in a discipline relevant to a specific Fellows workstream (e.g. economics, social sciences, or cybersecurity)
- * Experience in areas of research or engineering related to their workstream
- If you can't commit to the full duration, please still apply and note your constraints in the application
Qualifications
Must Haves
- * Are motivated by making sure AI is safe and beneficial for society as a whole
- * Are excited to transition into empirical AI research and would be interested in a full-time role at Anthropic
- * Have a strong technical background in computer science, mathematics, or physics
- * Thrive in fast-paced, collaborative environments
- * Can implement ideas quickly and communicate clearly
- * Fluent in Python programming
- * Available to work full-time on the Fellows program
- * Are motivated by reducing catastrophic risks from advanced AI systems
- * Have experience with empirical ML research projects
- * Have experience working with large language models
- * Have experience in one of the research areas mentioned above
- * Have a track record of open-source contributions
- * Are motivated by reducing catastrophic risks from advanced AI systems
- * Have contributed to open-source projects in LLM- or security-adjacent repositories
- * Have demonstrated success in bringing clarity and ownership to ambiguous technical problems
- * Have experience with pentesting, vulnerability research, or other offensive security work
- * Have a demonstrated willingness to do the "dirty work" that produces high-quality outputs
- * Have reported CVEs or been awarded bug bounties
- * Have experience with empirical ML research projects
- * Have experience with deep learning frameworks and experiment management
- To participate in the Fellows program, you must have work authorization in the US, UK, or Canada and be located in that country during the program
- We are also open to remote fellows in the UK, US, or Canada
- To participate in the Fellows program, you need to have or independently obtain full-time work authorization in the UK, the US, or Canada
- The program runs for 4 months, full-time
Nice to Haves
- * Strong background in a discipline relevant to a specific Fellows workstream (e.g. economics, social sciences, or cybersecurity)
- * Experience in areas of research or engineering related to their workstream
- If you can't commit to the full duration, please still apply and note your constraints in the application
Benefits
- Direct mentorship from Anthropic researchers
- Access to a shared workspace (in either Berkeley, California or London, UK)
- Connection to the broader AI safety and security research community
- Funding for compute (~$15k/month) and other research expenses
- We are also open to remote fellows in the UK, US, or Canada