tools
Ursa
Ursa is Speechmatics’ speech-to-text model family for transcribing live and recorded audio across languages, accents, and speakers.

In inglese
Ursa converts speech into text for applications such as captions, search, analytics, voice agents, legal transcription, and healthcare documentation. It supports batch and real-time transcription, speaker information, timestamps, formatting, and translation workflows.
Speechmatics released Ursa 2 as its next-generation model in October 2024. Current Speechmatics pricing is usage-based and is listed by API model and processing mode rather than by the Ursa name; translation, summaries, chapters, sentiment, and topics cost extra.
Features
- Batch transcription for recorded audio and video
- Real-time transcription with latency under one second
- Speech recognition across 55+ languages
- Speaker diarization for multi-speaker conversations
- Custom dictionaries and vocabulary support
- Word-level timestamps and subtitle formatting
- Cloud, on-premises, container, and on-device deployment
- Translation, summaries, sentiment, topics, and audio events
Use cases
- Generate live captions for broadcasts, meetings, and events
- Transcribe calls for contact-center analytics
- Build multilingual voice agents with speech-to-text APIs
- Create clinical, legal, or financial transcripts
- Search and index large audio and video archives
- Produce translated subtitles and accessible media
Pros
Cons
Capabilities
- Text to speech — “Text-to-Speech” source
- Speech to text — “Real-time speech-to-text is here” source
- Dubbing — “Our AI model supports 55+ languages for transcription, with 69 pairs supported for AI translation.” source
- Many languages — “Transcribe 55+ languages with a single model, even when speakers switch between languages naturally throughout a conversation.” source
- API — “Speech-to-text API built for every language and accent” source
Get it
Pricing
- Starting price
- Free
- Prices checked
- 2026-09-24
Free
Free
- No credit card required
- $100 credit to get started
- For developers and early exploration
- Speech-to-Text: 55+ languages
- 2 concurrent real-time sessions
- Text-to-Speech
Pro
- from $0.129/hr
- $0.129/hr
- $0.24/hr
- $0.40/hr
- $0.24/hr
- $0.43/hr
- $0.16/hr
- $0.011/1k characters
- 55+ languages
- 50 concurrent real-time sessions
- 10 file jobs per second
- Multi-region cloud options
- Low-latency Text-to-Speech
- Online email support
Enterprise
Price on request
- All our features, including audio alignment
- No rate limits
- Privacy-first deployment options
- Custom models
- SaaS or On-premises deployment
- Prioritized service and support