Vai al contenuto
AI.info

tools

Deepgram

Deepgram provides APIs for speech-to-text, text-to-speech, voice agents, and audio analysis.

Deepgram

In inglese

Deepgram is a developer platform for processing speech and building voice applications. Its APIs support real-time and batch transcription, speech synthesis, conversational voice agents, and audio analysis.

Developers, product teams, contact centers, and enterprises use it for transcription, voice interfaces, call analysis, and conversational systems. Usage is billed by audio, characters, tokens, or agent minutes; enterprise deployments and custom models require contacting sales.

Features

  • Real-time and pre-recorded speech-to-text APIs
  • Text-to-speech with Flux TTS, Aura-2, and Aura-1
  • Voice Agent API with turn detection and interruption handling
  • Audio Intelligence for summarization and conversation analysis
  • Speaker diarization, redaction, keyterm prompting, and smart formatting
  • Custom speech-to-text models for proprietary datasets
  • Self-hosted and private-cloud deployment options

Use cases

  • Transcribe live calls, meetings, podcasts, and recorded audio
  • Build conversational voice assistants and customer-service agents
  • Generate spoken responses for voice applications
  • Analyze conversations for summaries, entities, sentiment, and intent
  • Add speech recognition to contact-center and business workflows

Pros

    Cons

      Latest updates

      Capabilities

      • Text to speech — “Text-to-speech built for live conversations.” source
      • Speech to text — “Real-time transcription, production accuracy” source
      • Many languages — “Nova models support 45+ languages” source
      • API — “Deepgram unifies speech-to-text, text-to-speech, and LLM orchestration into a single API, reducing complexity, latency, and cost.” source

      Get it

      Pricing

      Starting price
      $4000/yr
      Prices checked
      2026-09-26
      Read
      by web search

      Pay As You Go

      Free

      • Free $200 Credit then pay-as-you-go
      • All endpoints in public models
      • Speech-to-Text: Up to 50 REST / 150 WSS / 5 Whisper Cloud
      • Text-to-Speech: Up to 45 REST + WSS
      • Voice Agent: Up to 45 WSS
      • Audio Intelligence: Up to 10 REST

      Growth

      • $4K+ / year
      • Save up to 20% with pre-paid credits for the year
      • All endpoints in public models
      • Speech-to-Text: Up to 50 REST / 225 WSS / 5 Whisper Cloud
      • Text-to-Speech: Up to 60 REST + WSS
      • Voice Agent: Up to 60 WSS
      • Audio Intelligence: Up to 10 REST

      Enterprise

      Price on request

      • For businesses with large volumes, data or deployment requirements, or support needs
      Official website