Summary
Epoch AI is a research institute that investigates trends in machine learning and the economic consequences of AI. They are seeking a Software Engineer to help evaluate frontier AI models, maintain benchmarking infrastructure, and contribute to the development of new benchmarks.
Responsibilities
- Implement benchmarks: Implement AI benchmarks within our evaluation infrastructure (primarily using the Inspect library) to expand the suite of capabilities we track. Develop our existing suite of benchmarks so we can quickly and painlessly evaluate new model releases
- Develop new benchmarks: Contribute to the development of brand new benchmarks. You will have the opportunity to pitch and prototype your own ideas in addition to helping out with existing projects
- Collaborate: Work closely with researchers, analysts, and other engineers at Epoch AI to ensure evaluation data and outputs are accurate, insightful, and effectively integrated into our research products and publications
Skills
- A strong software engineering background with more than two years of professional experience building and maintaining complex systems
- Ability to regularly contribute high-quality, robust, and maintainable code
- Comfortable diving deep into existing codebases and infrastructure
- Ability to generate ideas for new benchmarks, experiments, novel things to try, and other projects
- Motivated by Epoch AI's mission to provide rigorous, independent insight into key trends in AI
- Desire to deliver public, trustworthy evaluations of AI capabilities on challenging benchmarks
- Professional level English proficiency
- AI domain expertise is a strong plus but not required
- Hands-on experience running LLM evaluations
- Familiarity with evaluation frameworks like Inspect
- Solid grasp of current AI trends
- Ability to overlap with UTC–8 (Pacific Time) and UTC (Greenwich Mean Time)
- Willingness to travel for three retreats per year
Qualifications
Must Haves
- A strong software engineering background with more than two years of professional experience building and maintaining complex systems
- Ability to regularly contribute high-quality, robust, and maintainable code
- Comfortable diving deep into existing codebases and infrastructure
- Ability to generate ideas for new benchmarks, experiments, novel things to try, and other projects
- Motivated by Epoch AI's mission to provide rigorous, independent insight into key trends in AI
- Desire to deliver public, trustworthy evaluations of AI capabilities on challenging benchmarks
- Professional level English proficiency
Nice to Haves
- AI domain expertise is a strong plus but not required
- Hands-on experience running LLM evaluations
- Familiarity with evaluation frameworks like Inspect
- Solid grasp of current AI trends
- Ability to overlap with UTC–8 (Pacific Time) and UTC (Greenwich Mean Time)
- Willingness to travel for three retreats per year
Benefits
- Fully remote environment, including flexible work hours and schedules for most roles.
- Competitive global benefits program, including a comprehensive health insurance program—including supplemental benefits specific to a local country, as available and mandated by local law—and life insurance and a pension plan, if applicable in your country.
- Generous paid time off (PTO), including no specific limit on PTO with 30 days per year protected, unlimited personal and sick leave, and up to 6 months (combination of paid + unpaid) parental leave for permanent staff.
- A flexible and generous expense policy for you to spend on equipment and a large range of productivity tools or learning/development opportunities you might find valuable, subject to regulations and manager approval.
- Paid work trips, including 3 staff retreats per year and relevant conferences.
- Access to our very well-equipped offices in Berkeley, California, including paid meals, snacks, gym, and more. All staff, independently of where they are based, have access to the office for at least 20 days each year.