Summary
Gamma is focused on enhancing AI quality across its products, and they are seeking a Research Engineer to design evaluation frameworks and improve AI output quality. The role involves diagnosing AI-generated content failures, conducting experiments, and collaborating with product teams to ensure rapid and reliable AI improvements.
Responsibilities
- Design and maintain evaluation frameworks that measure AI output quality across all Gamma experiences, developing metrics and benchmarks to assess model performance
- Systematically improve production prompts through iterative experimentation, diagnosing failure patterns, crafting targeted improvements, and validating against quality benchmarks
- Conduct rigorous experiments to understand model behavior, analyze results, and derive insights that inform prompt and model improvements
- Build tools and workflows to support rapid experimentation and quality analysis, enabling faster iteration on AI improvements
- Fine-tune models on targeted datasets to improve baseline performance, preventing issues like poor layout choices or low-quality outlines
- Partner with product and engineering teams to ensure AI quality improvements ship quickly and work reliably at scale
Skills
- 2+ years working with AI systems, with demonstrated experience shipping production-grade AI products
- Deep hands-on experience with prompt engineering, LLM experimentation, and systematic evaluation of AI outputs
- Strong experimental mindset with the ability to design tests, analyze model performance, and iterate toward measurable quality improvements in ambiguous problem spaces
- Experience with post-training techniques for LLMs including reinforcement learning and supervised fine-tuning
- Exceptional attention to detail and genuine quality obsession, with care for output quality across all dimensions including less visible aspects
- Bachelor's degree in Computer Science, Machine Learning, or a related field, or equivalent hands-on experience with AI research and experimentation
Qualifications
Must Haves
- 2+ years working with AI systems, with demonstrated experience shipping production-grade AI products
- Deep hands-on experience with prompt engineering, LLM experimentation, and systematic evaluation of AI outputs
- Strong experimental mindset with the ability to design tests, analyze model performance, and iterate toward measurable quality improvements in ambiguous problem spaces
- Experience with post-training techniques for LLMs including reinforcement learning and supervised fine-tuning
- Exceptional attention to detail and genuine quality obsession, with care for output quality across all dimensions including less visible aspects
- Bachelor's degree in Computer Science, Machine Learning, or a related field, or equivalent hands-on experience with AI research and experimentation
Benefits