Summary
Gamma is focused on enhancing AI quality across its products, and they are seeking a Research Engineer to design evaluation frameworks and improve AI output quality. The role involves diagnosing failure patterns, conducting experiments, and collaborating with product and engineering teams to ensure effective AI implementations.
Responsibilities
- Design and maintain evaluation frameworks that measure AI output quality across all Gamma experiences, developing metrics and benchmarks to assess model performance
- Systematically improve production prompts through iterative experimentation, diagnosing failure patterns, crafting targeted improvements, and validating against quality benchmarks
- Conduct rigorous experiments to understand model behavior, analyze results, and derive insights that inform prompt and model improvements
- Build tools and workflows to support rapid experimentation and quality analysis, enabling faster iteration on AI improvements
- Fine-tune models on targeted datasets to improve baseline performance, preventing issues like poor layout choices or low-quality outlines
- Partner with product and engineering teams to ensure AI quality improvements ship quickly and work reliably at scale
Skills
- 2+ years working with AI systems, with demonstrated experience shipping production-grade AI products
- Deep hands-on experience with prompt engineering, LLM experimentation, and systematic evaluation of AI outputs
- Strong experimental mindset with the ability to design tests, analyze model performance, and iterate toward measurable quality improvements in ambiguous problem spaces
- Experience with post-training techniques for LLMs including reinforcement learning and supervised fine-tuning
- Exceptional attention to detail and genuine quality obsession, with care for output quality across all dimensions including less visible aspects
- Bachelor's degree in Computer Science, Machine Learning, or a related field, or equivalent hands-on experience with AI research and experimentation
Qualifications
Must Haves
- 2+ years working with AI systems, with demonstrated experience shipping production-grade AI products
- Deep hands-on experience with prompt engineering, LLM experimentation, and systematic evaluation of AI outputs
- Strong experimental mindset with the ability to design tests, analyze model performance, and iterate toward measurable quality improvements in ambiguous problem spaces
- Experience with post-training techniques for LLMs including reinforcement learning and supervised fine-tuning
- Exceptional attention to detail and genuine quality obsession, with care for output quality across all dimensions including less visible aspects
Nice to Haves
- Bachelor's degree in Computer Science, Machine Learning, or a related field, or equivalent hands-on experience with AI research and experimentation
Benefits