Skip to content
AI.info

Tools

AssemblyAI vs Ursa

AssemblyAI or Ursa? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.

In short

  • Both: Text to speech, Speech to text, Many languages, API
  • Only Ursa states: Dubbing

AssemblyAI

AssemblyAI provides APIs for transcribing, understanding, and building applications with spoken audio.

Plans

Free tier

Free

  • Up to 185 hours of pre-recorded transcription
  • Up to 333 hours of streaming transcription

Pre-recorded Speech-to-Text API — Universal-3.5 Pro

  • $0.21 /hr
  • Most accurate async speech-to-text model
  • Works across 18 languages
  • Native code switching

Pre-recorded Speech-to-Text API — Universal-2

  • $0.15 /hr
  • Supports 99 languages
  • Trained on over 12.5 million hours of audio data

Pre-recorded Speech-to-Text API — Custom

Price on request

  • Custom rate limits
  • Enhanced concurrency
  • Enterprise-grade flexibility

Realtime Speech-to-Text API — Universal-3.5 Pro Realtime

  • $0.45 /hr
  • Supports 18 languages
  • Context carryover and conversation memory
  • Self-correcting speaker labels and voice isolation

Realtime Speech-to-Text API — Universal-Streaming

  • $0.15 /hr
  • English-only transcription
  • Purpose-built for production voice applications

Realtime Speech-to-Text API — Universal-Streaming Multilingu

  • $0.15 /hr
  • Supports English, Spanish, German, French, Portuguese, and Italian

Realtime Speech-to-Text API — Custom

Price on request

  • Custom rate limits
  • Enhanced concurrency
  • Enterprise-grade flexibility

Sync Speech-to-Text API — Sync API

  • $0.45 /hr
  • Process up to 2 minutes per request
  • Results in ~134 ms (p50)
  • Per-word timing and confidence across 18 languages

Sync Speech-to-Text API — Custom

Price on request

  • Custom rate limits
  • Enhanced concurrency
  • Enterprise-grade flexibility

Prices checked 2026-09-25 on the maker’s page.

Capabilities

  • Text to speech — “Low-latency speech generation built for realtime conversation.” source
  • Speech to text — “Transcribes recorded audio and video asynchronously across 99 languages, with support on speaker diarization, PII redaction, and more features.” source
  • Many languages — “Transcribes recorded audio and video asynchronously across 99 languages, with support on speaker diarization, PII redaction, and more features.” source
  • API — “Build directly against the raw HTTP and WebSocket API for full control over requests and responses.” source

Security

  • SOC 2 Type II — “We maintain a SOC 2 Type 2 report covering security, availability, and confidentiality, audited annually by an independent firm.” source
  • SOC 2 — “We maintain a SOC 2 Type 2 report covering security, availability, and confidentiality, audited annually by an independent firm.” source
  • Not trained on your data — “Your data is never used to train or improve our models when you opt out.” source

Latest updates

About AssemblyAI

Ursa

Ursa is Speechmatics’ speech-to-text model family for transcribing live and recorded audio across languages, accents, and speakers.

Plans

Free

Free

  • No credit card required
  • $100 credit to get started
  • For developers and early exploration
  • Speech-to-Text: 55+ languages
  • 2 concurrent real-time sessions
  • Text-to-Speech

Pro

  • from $0.129/hr
  • $0.129/hr
  • $0.24/hr
  • $0.40/hr
  • $0.24/hr
  • $0.43/hr
  • $0.16/hr
  • $0.011/1k characters
  • 55+ languages
  • 50 concurrent real-time sessions
  • 10 file jobs per second
  • Multi-region cloud options
  • Low-latency Text-to-Speech
  • Online email support

Enterprise

Price on request

  • All our features, including audio alignment
  • No rate limits
  • Privacy-first deployment options
  • Custom models
  • SaaS or On-premises deployment
  • Prioritized service and support

Prices checked 2026-09-24 on the maker’s page.

Capabilities

  • Text to speech — “Text-to-Speech” source
  • Speech to text — “Real-time speech-to-text is here” source
  • Dubbing — “Our AI model supports 55+ languages for transcription, with 69 pairs supported for AI translation.” source
  • Many languages — “Transcribe 55+ languages with a single model, even when speakers switch between languages naturally throughout a conversation.” source
  • API — “Speech-to-text API built for every language and accent” source
About Ursa