Vai al contenuto
AI.info

tools

AssemblyAI

AssemblyAI provides APIs for transcribing, understanding, and building applications with spoken audio.

AssemblyAI

In inglese

AssemblyAI offers speech-to-text APIs for recorded and realtime audio, plus speaker diarization, language detection, prompting, sentiment analysis, entity detection, summaries, redaction, and other speech-understanding features.

Developers use it for voice agents, meeting notetakers, call analytics, medical documentation, dictation, and contact-center tools. It is an API platform rather than a consumer transcription app; usage is billed based on audio processed or tokens used, with some features costing extra.

Features

  • Transcribe recorded audio and video with Universal-3.5 Pro or Universal-2
  • Transcribe live audio with realtime speech-to-text APIs
  • Detect and label multiple speakers with speaker diarization
  • Improve recognition with keyterms prompting and natural-language prompting
  • Extract sentiment, entities, topics, translations, and summaries
  • Apply PII redaction, profanity filtering, and content moderation
  • Route transcripts to multiple language models through LLM Gateway
  • Run voice agents through a single WebSocket API

Use cases

  • Build realtime voice agents for customer support and sales
  • Transcribe meetings and generate summaries, chapters, and action items
  • Analyze customer calls for sentiment, topics, entities, and compliance
  • Create ambient medical documentation and AI scribe workflows
  • Add formatted speech input to dictation features
  • Provide live coaching and compliance alerts to contact-center agents

Pros

    Cons

      Latest updates

      Capabilities

      • Text to speech — “Low-latency speech generation built for realtime conversation.” source
      • Speech to text — “Transcribes recorded audio and video asynchronously across 99 languages, with support on speaker diarization, PII redaction, and more features.” source
      • Many languages — “Transcribes recorded audio and video asynchronously across 99 languages, with support on speaker diarization, PII redaction, and more features.” source
      • API — “Build directly against the raw HTTP and WebSocket API for full control over requests and responses.” source

      Get it

      Security

      • SOC 2 Type II — “We maintain a SOC 2 Type 2 report covering security, availability, and confidentiality, audited annually by an independent firm.” source
      • SOC 2 — “We maintain a SOC 2 Type 2 report covering security, availability, and confidentiality, audited annually by an independent firm.” source
      • Not trained on your data — “Your data is never used to train or improve our models when you opt out.” source

      Pricing

      Starting price
      Free
      Prices checked
      2026-09-25

      Free tier

      Free

      • Up to 185 hours of pre-recorded transcription
      • Up to 333 hours of streaming transcription

      Pre-recorded Speech-to-Text API — Universal-3.5 Pro

      • $0.21 /hr
      • Most accurate async speech-to-text model
      • Works across 18 languages
      • Native code switching

      Pre-recorded Speech-to-Text API — Universal-2

      • $0.15 /hr
      • Supports 99 languages
      • Trained on over 12.5 million hours of audio data

      Pre-recorded Speech-to-Text API — Custom

      Price on request

      • Custom rate limits
      • Enhanced concurrency
      • Enterprise-grade flexibility

      Realtime Speech-to-Text API — Universal-3.5 Pro Realtime

      • $0.45 /hr
      • Supports 18 languages
      • Context carryover and conversation memory
      • Self-correcting speaker labels and voice isolation

      Realtime Speech-to-Text API — Universal-Streaming

      • $0.15 /hr
      • English-only transcription
      • Purpose-built for production voice applications

      Realtime Speech-to-Text API — Universal-Streaming Multilingu

      • $0.15 /hr
      • Supports English, Spanish, German, French, Portuguese, and Italian

      Realtime Speech-to-Text API — Custom

      Price on request

      • Custom rate limits
      • Enhanced concurrency
      • Enterprise-grade flexibility

      Sync Speech-to-Text API — Sync API

      • $0.45 /hr
      • Process up to 2 minutes per request
      • Results in ~134 ms (p50)
      • Per-word timing and confidence across 18 languages

      Sync Speech-to-Text API — Custom

      Price on request

      • Custom rate limits
      • Enhanced concurrency
      • Enterprise-grade flexibility
      Official website