NewsBreak logo
NewsBreak
Posted 131 days agoVerified live 1d ago

Research Intern, Agent RL Training

Brief overview

Mountain View, CAIn-person
$35–$50/hrStated range
20 H-1B approvalsDept. of Labor
End-to-end model SFTPythonPyTorchMulti-node distributed training (FSDP, DeepSpeed, Megatron-LM)

About the company

NewsBreak logo
NewsBreaknewsbreak.com

NewsBreak is the leading platform that connects people with the information that truly matters, through technology

Visa sponsorship history

2 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
20H-1B approved
100%approval rate
13new H-1B hires
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20256
202614

Job description

Summary

NewsBreak is the Content Intelligence platform shaping the future content economy, with over 40 million monthly active users. The Research Intern will work closely with a mentor to explore the application of large language models to enhance NewsBreak’s core business functions.

Responsibilities

  • Collaborate with your full-time mentor to identify high-impact research directions for applying LLMs to NewsBreak’s products
  • Independently run end-to-end SFT experiments on LLM-based agents, and assist with RL-related exploration such as reward design and training iteration
  • Curate and build high-quality training datasets: instruction-following, preference pairs, agent trajectories, and synthetic data
  • Contribute to public publications; we encourage and support top-venue submissions during your internship

Skills

  • Highly motivated and committed: willing to put in extra hours when needed to push projects across the finish line
  • Genuine passion for research: you read papers for fun, tinker with models on weekends, and care deeply about advancing the field
  • Independently capable of end-to-end model SFT: with basic understanding of RL-based post-training methods (RLHF, DPO, PPO, GRPO, etc.)
  • Excellent taste in model behavior: able to reason about what 'good' looks like across user-facing domains and articulate why
  • Strong Python and PyTorch skills
  • Publication at a top-tier venue (NeurIPS, ICML, ICLR, ACL, EMNLP, or equivalent)
  • Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM)
  • Proficiency in writing custom GPU kernels with Triton or CUDA
  • Experience building synthetic data pipelines for agent training
  • Familiarity with open-source RL frameworks: TRL, OpenRLHF, veRL/vLLM

Qualifications

Must Haves

  • Highly motivated and committed: willing to put in extra hours when needed to push projects across the finish line
  • Genuine passion for research: you read papers for fun, tinker with models on weekends, and care deeply about advancing the field
  • Independently capable of end-to-end model SFT: with basic understanding of RL-based post-training methods (RLHF, DPO, PPO, GRPO, etc.)
  • Excellent taste in model behavior: able to reason about what 'good' looks like across user-facing domains and articulate why
  • Strong Python and PyTorch skills

Nice to Haves

  • Publication at a top-tier venue (NeurIPS, ICML, ICLR, ACL, EMNLP, or equivalent)
  • Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM)
  • Proficiency in writing custom GPU kernels with Triton or CUDA
  • Experience building synthetic data pipelines for agent training
  • Familiarity with open-source RL frameworks: TRL, OpenRLHF, veRL/vLLM

Benefits

  • Depending on the position, the role may also be eligible for discretionary bonus and options.

More jobs like this