tools
Ursa
Ursa is Speechmatics’ speech-to-text model family for transcribing live and recorded audio across languages, accents, and speakers.

Ursa converts speech into text for applications such as captions, search, analytics, voice agents, legal transcription, and healthcare documentation. It supports batch and real-time transcription, speaker information, timestamps, formatting, and translation workflows.
Speechmatics released Ursa 2 as its next-generation model in October 2024. Current Speechmatics pricing is usage-based and is listed by API model and processing mode rather than by the Ursa name; translation, summaries, chapters, sentiment, and topics cost extra.
Features
- Batch transcription for recorded audio and video
- Real-time transcription with latency under one second
- Speech recognition across 55+ languages
- Speaker diarization for multi-speaker conversations
- Custom dictionaries and vocabulary support
- Word-level timestamps and subtitle formatting
- Cloud, on-premises, container, and on-device deployment
- Translation, summaries, sentiment, topics, and audio events
Use cases
- Generate live captions for broadcasts, meetings, and events
- Transcribe calls for contact-center analytics
- Build multilingual voice agents with speech-to-text APIs
- Create clinical, legal, or financial transcripts
- Search and index large audio and video archives
- Produce translated subtitles and accessible media
Pros
Cons
Pricing
- Starting price
- Free
- Pricing checked
- 2026-09-19
Free
$100 free
- 55+ languages
- 2 concurrent real-time sessions
- Multi-region cloud options
- Text-to-speech
Pro
from$0.129/hr
- 55+ languages
- 50 concurrent real-time sessions
- 10 file jobs per second
- Multi-region cloud options
- Text-to-speech
Enterprise
Custom
- Unlimited scale
- Private cloud, container, virtual appliance, or on-device deployment
- Unlimited real-time session concurrency
- Unlimited batch job creation