Summary
Phonic is a product and research lab focused on powering realistic, human-like voice AI conversations. As a Research Intern, you'll own research directions end-to-end, from identifying problems to designing experiments and developing methods that lead to production impact.
Responsibilities
- Identify high-leverage research problems across the voice AI stack from audio understanding to audio output, and take full ownership of driving them forward
- Design and run rigorous experiments that analyze architectural trade-offs to understand how design choices influence a model鈥檚 scalability, latency, and quality
- Curate massive training datasets and execute rigorous experiments to determine exactly how data quality shapes model behavior and performance
- Work directly with research scientists and engineers to move fast from prototype to production
- Build the training pipelines, evaluation frameworks, and tooling that let us experiment and iterate quickly
Skills
- A track record of original work: you've found a real problem, developed an approach, and seen it through
- Proficiency in PyTorch (or JAX), and the ability to implement models cleanly from papers
- Fluency in the math, probability, optimization, and linear algebra underlying model behavior
- You move fluidly between ideas and implementation; you don't just think about problems, you build things
- Clear, precise written and verbal communication
- Research experience in speech, audio, or language modeling (ASR, TTS, LLMs, codec models)
- Familiarity with generative modeling techniques: diffusion, flow matching, or autoregressive models
- Experience with RLHF or preference optimization
- Competitive programming or olympiad background
- Publications or preprints at venues like NeurIPS, ICML, ICLR, Interspeech, ICASSP, or ACL
Qualifications
Must Haves
- A track record of original work: you've found a real problem, developed an approach, and seen it through
- Proficiency in PyTorch (or JAX), and the ability to implement models cleanly from papers
- Fluency in the math, probability, optimization, and linear algebra underlying model behavior
- You move fluidly between ideas and implementation; you don't just think about problems, you build things
- Clear, precise written and verbal communication
Nice to Haves
- Research experience in speech, audio, or language modeling (ASR, TTS, LLMs, codec models)
- Familiarity with generative modeling techniques: diffusion, flow matching, or autoregressive models
- Experience with RLHF or preference optimization
- Competitive programming or olympiad background
- Publications or preprints at venues like NeurIPS, ICML, ICLR, Interspeech, ICASSP, or ACL
Benefits
- 馃捀 Top-tier compensation: in order to get the best talent, we provide salary and equity that recognize your skillset
- 馃 Meals: free breakfast, lunch, and dinner provided in the office
- 馃 We have regular off-sites and team celebrations