Tools
AssemblyAI vs Deepgram
AssemblyAI or Deepgram? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In short
- Both: Text to speech, Speech to text, Many languages, API
AssemblyAI
AssemblyAI provides APIs for transcribing, understanding, and building applications with spoken audio.
Plans
Free tier
Free
- Up to 185 hours of pre-recorded transcription
- Up to 333 hours of streaming transcription
Pre-recorded Speech-to-Text API — Universal-3.5 Pro
- $0.21 /hr
- Most accurate async speech-to-text model
- Works across 18 languages
- Native code switching
Pre-recorded Speech-to-Text API — Universal-2
- $0.15 /hr
- Supports 99 languages
- Trained on over 12.5 million hours of audio data
Pre-recorded Speech-to-Text API — Custom
Price on request
- Custom rate limits
- Enhanced concurrency
- Enterprise-grade flexibility
Realtime Speech-to-Text API — Universal-3.5 Pro Realtime
- $0.45 /hr
- Supports 18 languages
- Context carryover and conversation memory
- Self-correcting speaker labels and voice isolation
Realtime Speech-to-Text API — Universal-Streaming
- $0.15 /hr
- English-only transcription
- Purpose-built for production voice applications
Realtime Speech-to-Text API — Universal-Streaming Multilingu
- $0.15 /hr
- Supports English, Spanish, German, French, Portuguese, and Italian
Realtime Speech-to-Text API — Custom
Price on request
- Custom rate limits
- Enhanced concurrency
- Enterprise-grade flexibility
Sync Speech-to-Text API — Sync API
- $0.45 /hr
- Process up to 2 minutes per request
- Results in ~134 ms (p50)
- Per-word timing and confidence across 18 languages
Sync Speech-to-Text API — Custom
Price on request
- Custom rate limits
- Enhanced concurrency
- Enterprise-grade flexibility
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Text to speech — “Low-latency speech generation built for realtime conversation.” source
- Speech to text — “Transcribes recorded audio and video asynchronously across 99 languages, with support on speaker diarization, PII redaction, and more features.” source
- Many languages — “Transcribes recorded audio and video asynchronously across 99 languages, with support on speaker diarization, PII redaction, and more features.” source
- API — “Build directly against the raw HTTP and WebSocket API for full control over requests and responses.” source
Security
- SOC 2 Type II — “We maintain a SOC 2 Type 2 report covering security, availability, and confidentiality, audited annually by an independent firm.” source
- SOC 2 — “We maintain a SOC 2 Type 2 report covering security, availability, and confidentiality, audited annually by an independent firm.” source
- Not trained on your data — “Your data is never used to train or improve our models when you opt out.” source
Latest updates
- New LLM Gateway Models: NVIDIA Nemotron, DeepSeek v4.1 Flash, GLM 5.3, & GLM 5.3 Flash
Four new models are now live in LLM Gateway: NVIDIA's Nemotron models, DeepSeek's v4.1 Flash, and Zhipu AI's GLM 5.3 and GLM 5.3 Flash.
- Four NVIDIA Nemotron Models Now Available on the LLM Gateway
Four new NVIDIA models are now live in LLM Gateway: Nemotron 3 Nano 30B A3B, Nemotron 3 Super 120B A12B, Nemotron Lightning 3.5 30B A3B, and Nemotron Nano 9B v2.
- Gemini 3.8 Flash Now Available on the LLM Gateway
Google's gemini-3.8-flash is now available through the LLM Gateway. It has a 1M-token context window and supports tool calling, structured outputs, and streaming.
Deepgram
Deepgram provides APIs for speech-to-text, text-to-speech, voice agents, and audio analysis.
Plans
Pay As You Go
Free
- Free $200 Credit then pay-as-you-go
- All endpoints in public models
- Speech-to-Text: Up to 50 REST / 150 WSS / 5 Whisper Cloud
- Text-to-Speech: Up to 45 REST + WSS
- Voice Agent: Up to 45 WSS
- Audio Intelligence: Up to 10 REST
Growth
- $4K+ / year
- Save up to 20% with pre-paid credits for the year
- All endpoints in public models
- Speech-to-Text: Up to 50 REST / 225 WSS / 5 Whisper Cloud
- Text-to-Speech: Up to 60 REST + WSS
- Voice Agent: Up to 60 WSS
- Audio Intelligence: Up to 10 REST
Enterprise
Price on request
- For businesses with large volumes, data or deployment requirements, or support needs
Prices checked 2026-09-26 by web search.
Capabilities
- Text to speech — “Text-to-speech built for live conversations.” source
- Speech to text — “Real-time transcription, production accuracy” source
- Many languages — “Nova models support 45+ languages” source
- API — “Deepgram unifies speech-to-text, text-to-speech, and LLM orchestration into a single API, reducing complexity, latency, and cost.” source
Latest updates
- Nova-3 Improved Models for Flemish, German (Switzerland), Lithuanian, and Portuguese
Released improved Nova-3 monolingual models for four existing languages, enhancing transcription quality for batch and streaming workloads.
- Nova-3 Improved Models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu, and Vietnamese
Released improved Nova-3 monolingual models for nine existing languages, enhancing transcription quality for batch and streaming workloads.
- Nova-3 Pharma: Speech-to-Text for Pharmaceutical Use Cases (English)
Released nova-3-pharma, a new Nova-3 model purpose-built for pharmaceutical vocabulary, with a focus on accurate drug-name recognition.