tools
Deepgram
Deepgram provides APIs for speech-to-text, text-to-speech, voice agents, and audio analysis.

In inglese
Deepgram is a developer platform for processing speech and building voice applications. Its APIs support real-time and batch transcription, speech synthesis, conversational voice agents, and audio analysis.
Developers, product teams, contact centers, and enterprises use it for transcription, voice interfaces, call analysis, and conversational systems. Usage is billed by audio, characters, tokens, or agent minutes; enterprise deployments and custom models require contacting sales.
Features
- Real-time and pre-recorded speech-to-text APIs
- Text-to-speech with Flux TTS, Aura-2, and Aura-1
- Voice Agent API with turn detection and interruption handling
- Audio Intelligence for summarization and conversation analysis
- Speaker diarization, redaction, keyterm prompting, and smart formatting
- Custom speech-to-text models for proprietary datasets
- Self-hosted and private-cloud deployment options
Use cases
- Transcribe live calls, meetings, podcasts, and recorded audio
- Build conversational voice assistants and customer-service agents
- Generate spoken responses for voice applications
- Analyze conversations for summaries, entities, sentiment, and intent
- Add speech recognition to contact-center and business workflows
Pros
Cons
Latest updates
- Nova-3 Improved Models for Flemish, German (Switzerland), Lithuanian, and Portuguese
Released improved Nova-3 monolingual models for four existing languages, enhancing transcription quality for batch and streaming workloads.
- Nova-3 Improved Models for Danish, Estonian, Flemish, Italian, Lithuanian, Macedonian, Polish, Urdu, and Vietnamese
Released improved Nova-3 monolingual models for nine existing languages, enhancing transcription quality for batch and streaming workloads.
- Nova-3 Pharma: Speech-to-Text for Pharmaceutical Use Cases (English)
Released nova-3-pharma, a new Nova-3 model purpose-built for pharmaceutical vocabulary, with a focus on accurate drug-name recognition.
- India Endpoint Now Generally Available
The Deepgram India endpoint (api.in.deepgram.com) is now generally available for customers requiring data processing within India.
- Deepgram Self-Hosted September 2026 Release (260915) (260915)
Self-hosted release 260915; minimum required NVIDIA driver version is >=580, open kernel module flavor.
Capabilities
- Text to speech — “Text-to-speech built for live conversations.” source
- Speech to text — “Real-time transcription, production accuracy” source
- Many languages — “Nova models support 45+ languages” source
- API — “Deepgram unifies speech-to-text, text-to-speech, and LLM orchestration into a single API, reducing complexity, latency, and cost.” source
Get it
Pricing
- Starting price
- $4000/yr
- Prices checked
- 2026-09-26
- Read
- by web search
Pay As You Go
Free
- Free $200 Credit then pay-as-you-go
- All endpoints in public models
- Speech-to-Text: Up to 50 REST / 150 WSS / 5 Whisper Cloud
- Text-to-Speech: Up to 45 REST + WSS
- Voice Agent: Up to 45 WSS
- Audio Intelligence: Up to 10 REST
Growth
- $4K+ / year
- Save up to 20% with pre-paid credits for the year
- All endpoints in public models
- Speech-to-Text: Up to 50 REST / 225 WSS / 5 Whisper Cloud
- Text-to-Speech: Up to 60 REST + WSS
- Voice Agent: Up to 60 WSS
- Audio Intelligence: Up to 10 REST
Enterprise
Price on request
- For businesses with large volumes, data or deployment requirements, or support needs