Tools
Bulbul vs Saaras
Bulbul or Saaras? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In inglese
In short
- Both: API
- Only Bulbul states: Text to speech, Voice cloning, Commercial use
- Only Saaras states: Speech to text, Many languages
Bulbul
Bulbul is Sarvam's text-to-speech model for generating speech in 11 Indian languages with selectable voices and pace control.
Plans
Bulbul v3
- ₹30 per 10K characters
- Rounded to the nearest character.
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Text to speech — “Bulbul v3 is our latest text-to-speech model, specifically designed for Indian languages and accents.” source
- Voice cloning — “Bulbul V3 supports voice cloning, allowing teams to create custom voices that maintain natural expressiveness and quality.” source
- Commercial use — “For commercial production rights when shipping generated audio, see Commercial Licensing.” source
- API — “Looking to integrate the Bulbul V3 API within your Products/Applications?” source
Security
- SOC 2 Type II — “ISO 27001 and SOC 2 Type II certified.” source
- SOC 2 — “ISO 27001 and SOC 2 Type II certified.” source
- ISO 27001 — “We're ISO 27001:2022 certified and hold SOC 2 Type II.” source
- Not: ISO 42001 — “ISO 42001 is in progress, and we'll publish it when it's done.” source
- Not: Not trained on your data — “Custom models trained on a customer's data remain inside the customer's environment, with weights they own and we never reuse.” source
Latest updates
- Sarvam Vision 2.1: Pushing the Pareto frontier of document intelligence (2.1)
Sarvam Vision 2.1 adds structured extraction and Indic handwritten recognition capabilities.
- Introducing Saaras V4 (V4)
Introduces Saaras V4, an ASR model for a multilingual world.
- Everything we announced at Sarvam Epoch
Announces new models, Sarvam Inference, Indus agents, Sarvam Code, and more.
Saaras
Saaras is Sarvam AI's speech-to-text model for transcribing, translating, and processing speech in Indian languages.
Plans
Starter (Pay as you go)
- ₹3.00 per 1,000 characters (Text to speech, Real-time)
- ₹3.00 per 1,000 characters (Text to speech, Streaming)
- ₹30.00 per hour (Speech to text, Real-time)
- ₹30.00 per hour (Speech to text, Streaming)
- ₹30.00 per hour (Speech to text, Batch)
- ₹45.00 per hour (Speech to text, Batch with diarization)
- ₹128.10 per 1M tokens (GLM 5.2, input)
- ₹23.79 per 1M tokens (GLM 5.2, cached input)
- ₹402.60 per 1M tokens (GLM 5.2, output)
- ₹36.60 per 1M tokens (Gemma-4 31B, input)
- ₹13.73 per 1M tokens (Gemma-4 31B, cached input)
- ₹91.50 per 1M tokens (Gemma-4 31B, output)
- ₹29.28 per 1M tokens (Sarvam 105B, input)
- ₹10.98 per 1M tokens (Sarvam 105B, cached input)
- ₹73.20 per 1M tokens (Sarvam 105B, output)
- ₹29.28 per 1M tokens (Sarvam 105B Chat, input)
- ₹10.98 per 1M tokens (Sarvam 105B Chat, cached input)
- ₹73.20 per 1M tokens (Sarvam 105B Chat, output)
- ₹29.28 per 1M tokens (Sarvam 105B Conversations, input)
- ₹10.98 per 1M tokens (Sarvam 105B Conversations, cached input)
- ₹73.20 per 1M tokens (Sarvam 105B Conversations, output)
- ₹0.005 per character (Translate & transliterate, Pay as you go)
- ₹0.50 per page (Digitisation API)
- ₹1.00 per page (Extraction API)
- ₹1.67 per second (Dubbing API, Pay as you go)
- ₹0.005 (Translate & transliterate, Starter)
- ₹1.67 (Dubbing API, Starter)
- Start with free credits, then add exactly what you need
- Universal credits across supported APIs
- Credits never expire
- No monthly commitment
Enterprise
Price on request
- Tailored limits, throughput, pricing, and support
- Higher rate limits on every API
- Custom concurrency & throughput
- Volume discounts on credits
- Dedicated support & SLAs
Pro
- ₹0.005 per character (Translate & transliterate, Pro)
- ₹1.67 per second (Dubbing API, Pro)
Growth
- ₹0.0045 per character (Translate & transliterate, Growth)
- ₹1.25 per second (Dubbing API, Growth)
Business
- ₹0.004 per character (Translate & transliterate, Business)
- ₹1.20 per second (Dubbing API, Business)
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Speech to text — “From the same audio input, it can produce five transcript modes: verbatim, normalized, code-mixed, transliterated, and translated.” source
- Many languages — “Saaras V4 achieved SOTA performance on all 22 Indian languages and reaffirms our commitment to even the most low-resource Indian languages” source
- API — “Integrate Saaras v4 into your application with a simple API key from the Sarvam AI dashboard” source
Security
- SOC 2 Type II — “Complete data residency in India. ISO 27001 and SOC 2 Type II certified.” source
- SOC 2 — “Complete data residency in India. ISO 27001 and SOC 2 Type II certified.” source
- ISO 27001 — “We're ISO 27001:2022 certified and hold SOC 2 Type II.” source
- Not trained on your data — “Customer data is never used to train models for other customers” source
Latest updates
- Sarvam Vision 2.1: Pushing the Pareto frontier of document intelligence (2.1)
Sarvam Vision 2.1 adds structured extraction and Indic handwritten recognition.
- Introducing Saaras V4 (V4)
Introduced Saaras V4 for multilingual ASR.
- Everything we announced at Sarvam Epoch
Announced new models, Sarvam Inference, Indus agents, Sarvam Code, and more.