Tools
Deepgram vs Ursa
Deepgram or Ursa? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In short
- Both: Text to speech, Speech to text, Many languages, API
- Only Ursa states: Dubbing
Deepgram
Deepgram provides APIs for speech-to-text, text-to-speech, voice agents, and audio analysis.
Plans
Pay As You Go
Free
- Free $200 Credit then pay-as-you-go
- All endpoints in public models
- Speech-to-Text: Up to 50 REST / 150 WSS / 5 Whisper Cloud
- Text-to-Speech: Up to 45 REST + WSS
- Voice Agent: Up to 45 WSS
- Audio Intelligence: Up to 10 REST
Growth
- $4K+ / year
- Save up to 20% with pre-paid credits for the year
- All endpoints in public models
- Speech-to-Text: Up to 50 REST / 225 WSS / 5 Whisper Cloud
- Text-to-Speech: Up to 60 REST + WSS
- Voice Agent: Up to 60 WSS
- Audio Intelligence: Up to 10 REST
Enterprise
Price on request
- For businesses with large volumes, data or deployment requirements, or support needs
Prices checked 2026-09-26 by web search.
Capabilities
- Text to speech — “Text-to-speech built for live conversations.” source
- Speech to text — “Real-time transcription, production accuracy” source
- Many languages — “Nova models support 45+ languages” source
- API — “Deepgram unifies speech-to-text, text-to-speech, and LLM orchestration into a single API, reducing complexity, latency, and cost.” source
Latest updates
- Nova-3 Improved Models for Flemish, German (Switzerland), Lithuanian, and Portuguese
Released improved Nova-3 monolingual models for four existing languages, enhancing transcription quality for batch and streaming workloads.
- Nova-3 Improved Models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu, and Vietnamese
Released improved Nova-3 monolingual models for nine existing languages, enhancing transcription quality for batch and streaming workloads.
- Nova-3 Pharma: Speech-to-Text for Pharmaceutical Use Cases (English)
Released nova-3-pharma, a new Nova-3 model purpose-built for pharmaceutical vocabulary, with a focus on accurate drug-name recognition.
Ursa
Ursa is Speechmatics’ speech-to-text model family for transcribing live and recorded audio across languages, accents, and speakers.
Plans
Free
Free
- No credit card required
- $100 credit to get started
- For developers and early exploration
- Speech-to-Text: 55+ languages
- 2 concurrent real-time sessions
- Text-to-Speech
Pro
- from $0.129/hr
- $0.129/hr
- $0.24/hr
- $0.40/hr
- $0.24/hr
- $0.43/hr
- $0.16/hr
- $0.011/1k characters
- 55+ languages
- 50 concurrent real-time sessions
- 10 file jobs per second
- Multi-region cloud options
- Low-latency Text-to-Speech
- Online email support
Enterprise
Price on request
- All our features, including audio alignment
- No rate limits
- Privacy-first deployment options
- Custom models
- SaaS or On-premises deployment
- Prioritized service and support
Prices checked 2026-09-24 on the maker’s page.
Capabilities
- Text to speech — “Text-to-Speech” source
- Speech to text — “Real-time speech-to-text is here” source
- Dubbing — “Our AI model supports 55+ languages for transcription, with 69 pairs supported for AI translation.” source
- Many languages — “Transcribe 55+ languages with a single model, even when speakers switch between languages naturally throughout a conversation.” source
- API — “Speech-to-text API built for every language and accent” source