jobs
Staff Research Engineer - Multimodal Generative Modelling
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way
- Company
- Synthesia
- Location
- Europe
- Status
- Open
- Posted
- 2026-07-17T14:38:23.387+00:00
Synthesia is hiring a staff research engineer for its Voice team, part of a 40-plus-person R&D group, to build multimodal generative models that combine text, audio and video. The work covers roadmap definition, novel architecture proposals, pretraining through post-training (DPO, fine-tuning, distillation), evaluation metrics for latency-aware conversation, dataset curation and shipping optimised models to production. The posting asks for hands-on LLM or transformer experience, distributed PyTorch, time-series modelling and tokenisation in audio or video, and fast prototyping. Scope spans voice, video and product.
Original job posting