Tools
AssemblyAI vs Saaras
AssemblyAI or Saaras? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In inglese
In short
- Both: Speech to text, Many languages, API
- Only AssemblyAI states: Text to speech
AssemblyAI
AssemblyAI provides APIs for transcribing, understanding, and building applications with spoken audio.
Plans
Free tier
Free
- Up to 185 hours of pre-recorded transcription
- Up to 333 hours of streaming transcription
Pre-recorded Speech-to-Text API — Universal-3.5 Pro
- $0.21 /hr
- Most accurate async speech-to-text model
- Works across 18 languages
- Native code switching
Pre-recorded Speech-to-Text API — Universal-2
- $0.15 /hr
- Supports 99 languages
- Trained on over 12.5 million hours of audio data
Pre-recorded Speech-to-Text API — Custom
Price on request
- Custom rate limits
- Enhanced concurrency
- Enterprise-grade flexibility
Realtime Speech-to-Text API — Universal-3.5 Pro Realtime
- $0.45 /hr
- Supports 18 languages
- Context carryover and conversation memory
- Self-correcting speaker labels and voice isolation
Realtime Speech-to-Text API — Universal-Streaming
- $0.15 /hr
- English-only transcription
- Purpose-built for production voice applications
Realtime Speech-to-Text API — Universal-Streaming Multilingu
- $0.15 /hr
- Supports English, Spanish, German, French, Portuguese, and Italian
Realtime Speech-to-Text API — Custom
Price on request
- Custom rate limits
- Enhanced concurrency
- Enterprise-grade flexibility
Sync Speech-to-Text API — Sync API
- $0.45 /hr
- Process up to 2 minutes per request
- Results in ~134 ms (p50)
- Per-word timing and confidence across 18 languages
Sync Speech-to-Text API — Custom
Price on request
- Custom rate limits
- Enhanced concurrency
- Enterprise-grade flexibility
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Text to speech — “Low-latency speech generation built for realtime conversation.” source
- Speech to text — “Transcribes recorded audio and video asynchronously across 99 languages, with support on speaker diarization, PII redaction, and more features.” source
- Many languages — “Transcribes recorded audio and video asynchronously across 99 languages, with support on speaker diarization, PII redaction, and more features.” source
- API — “Build directly against the raw HTTP and WebSocket API for full control over requests and responses.” source
Security
- SOC 2 Type II — “We maintain a SOC 2 Type 2 report covering security, availability, and confidentiality, audited annually by an independent firm.” source
- SOC 2 — “We maintain a SOC 2 Type 2 report covering security, availability, and confidentiality, audited annually by an independent firm.” source
- Not trained on your data — “Your data is never used to train or improve our models when you opt out.” source
Latest updates
- New LLM Gateway Models: NVIDIA Nemotron, DeepSeek v4.1 Flash, GLM 5.3, & GLM 5.3 Flash
Four new models are now live in LLM Gateway: NVIDIA's Nemotron models, DeepSeek's v4.1 Flash, and Zhipu AI's GLM 5.3 and GLM 5.3 Flash.
- Four NVIDIA Nemotron Models Now Available on the LLM Gateway
Four new NVIDIA models are now live in LLM Gateway: Nemotron 3 Nano 30B A3B, Nemotron 3 Super 120B A12B, Nemotron Lightning 3.5 30B A3B, and Nemotron Nano 9B v2.
- Gemini 3.8 Flash Now Available on the LLM Gateway
Google's gemini-3.8-flash is now available through the LLM Gateway. It has a 1M-token context window and supports tool calling, structured outputs, and streaming.
Saaras
Saaras is Sarvam AI's speech-to-text model for transcribing, translating, and processing speech in Indian languages.
Plans
Starter (Pay as you go)
- ₹3.00 per 1,000 characters (Text to speech, Real-time)
- ₹3.00 per 1,000 characters (Text to speech, Streaming)
- ₹30.00 per hour (Speech to text, Real-time)
- ₹30.00 per hour (Speech to text, Streaming)
- ₹30.00 per hour (Speech to text, Batch)
- ₹45.00 per hour (Speech to text, Batch with diarization)
- ₹128.10 per 1M tokens (GLM 5.2, input)
- ₹23.79 per 1M tokens (GLM 5.2, cached input)
- ₹402.60 per 1M tokens (GLM 5.2, output)
- ₹36.60 per 1M tokens (Gemma-4 31B, input)
- ₹13.73 per 1M tokens (Gemma-4 31B, cached input)
- ₹91.50 per 1M tokens (Gemma-4 31B, output)
- ₹29.28 per 1M tokens (Sarvam 105B, input)
- ₹10.98 per 1M tokens (Sarvam 105B, cached input)
- ₹73.20 per 1M tokens (Sarvam 105B, output)
- ₹29.28 per 1M tokens (Sarvam 105B Chat, input)
- ₹10.98 per 1M tokens (Sarvam 105B Chat, cached input)
- ₹73.20 per 1M tokens (Sarvam 105B Chat, output)
- ₹29.28 per 1M tokens (Sarvam 105B Conversations, input)
- ₹10.98 per 1M tokens (Sarvam 105B Conversations, cached input)
- ₹73.20 per 1M tokens (Sarvam 105B Conversations, output)
- ₹0.005 per character (Translate & transliterate, Pay as you go)
- ₹0.50 per page (Digitisation API)
- ₹1.00 per page (Extraction API)
- ₹1.67 per second (Dubbing API, Pay as you go)
- ₹0.005 (Translate & transliterate, Starter)
- ₹1.67 (Dubbing API, Starter)
- Start with free credits, then add exactly what you need
- Universal credits across supported APIs
- Credits never expire
- No monthly commitment
Enterprise
Price on request
- Tailored limits, throughput, pricing, and support
- Higher rate limits on every API
- Custom concurrency & throughput
- Volume discounts on credits
- Dedicated support & SLAs
Pro
- ₹0.005 per character (Translate & transliterate, Pro)
- ₹1.67 per second (Dubbing API, Pro)
Growth
- ₹0.0045 per character (Translate & transliterate, Growth)
- ₹1.25 per second (Dubbing API, Growth)
Business
- ₹0.004 per character (Translate & transliterate, Business)
- ₹1.20 per second (Dubbing API, Business)
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Speech to text — “From the same audio input, it can produce five transcript modes: verbatim, normalized, code-mixed, transliterated, and translated.” source
- Many languages — “Saaras V4 achieved SOTA performance on all 22 Indian languages and reaffirms our commitment to even the most low-resource Indian languages” source
- API — “Integrate Saaras v4 into your application with a simple API key from the Sarvam AI dashboard” source
Security
- SOC 2 Type II — “Complete data residency in India. ISO 27001 and SOC 2 Type II certified.” source
- SOC 2 — “Complete data residency in India. ISO 27001 and SOC 2 Type II certified.” source
- ISO 27001 — “We're ISO 27001:2022 certified and hold SOC 2 Type II.” source
- Not trained on your data — “Customer data is never used to train models for other customers” source
Latest updates
- Sarvam Vision 2.1: Pushing the Pareto frontier of document intelligence (2.1)
Sarvam Vision 2.1 adds structured extraction and Indic handwritten recognition.
- Introducing Saaras V4 (V4)
Introduced Saaras V4 for multilingual ASR.
- Everything we announced at Sarvam Epoch
Announced new models, Sarvam Inference, Indus agents, Sarvam Code, and more.