jobs
Research Intern, Inference (Summer 2027)
About The Role The Inference Research team is dedicated to building the next generation of efficient, scalable, and reliable serving systems for large foundation models, directly contributing to the mission of advancing open and transparent
In inglese
- Company
- Together AI
- Location
- San Francisco
- Status
- Open
- Posted
- 2026-09-18T15:44:40+00:00
A 12-to-14-week research internship on Together AI's Inference Research team in San Francisco, working on serving systems for large foundation models. Projects span distributed inference, cross-layer optimization across models, systems and hardware, KV cache design, speculative decoding and large-scale serving. The intern runs experiments, reports progress and writes up findings in papers and blog posts. Applicants should be finishing a Bachelor's, Master's or PhD in computer science, electrical engineering or a related field, with PyTorch or JAX and strong Python; CUDA experience and MLSys or ICLR publications are preferred.
Original job posting