Description
NewsBreak is hiring a Research Intern for its Agent RL Training team in Mountain View, California. The intern will work with a full-time mentor to explore applying large language models to content understanding, recommendations, agentic web browsing, and autonomous multi-step task completion. Responsibilities include independently running end-to-end supervised fine-tuning experiments, exploring reward design and RL training iteration, curating training datasets, and contributing to publications. The role requires strong Python and PyTorch skills, basic understanding of RL-based post-training methods, and the ability to reason about model behavior; preferred qualifications include top-tier publication experience, distributed training, GPU kernel development, synthetic data pipelines, and open-source RL frameworks. The position pays $35–$50 USD per hour.
