Description
Together AI is hiring a Research Intern to work on distributed inference, compiler-aware optimization, speculative decoding, phase-aware execution, KV cache design, and large-scale serving architectures for foundation models. The intern will design experiments, communicate project progress, and document findings in scientific publications and blog posts. The internship runs for 12 to 14 weeks, with cohorts from May 17 to August 6 or June 14 to September 3, and offers an estimated US hourly rate of $58 to $70 plus housing stipends and other benefits.
