Tools
Bulbul vs Coqui XTTS
Bulbul or Coqui XTTS? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In short
- Both: Text to speech, Voice cloning, API
- Only Bulbul states: Commercial use
- Only Coqui XTTS states: Many languages
Bulbul
Bulbul is Sarvam's text-to-speech model for generating speech in 11 Indian languages with selectable voices and pace control.
Plans
Bulbul v3
- ₹30 per 10K characters
- Rounded to the nearest character.
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Text to speech — “Bulbul v3 is our latest text-to-speech model, specifically designed for Indian languages and accents.” source
- Voice cloning — “Bulbul V3 supports voice cloning, allowing teams to create custom voices that maintain natural expressiveness and quality.” source
- Commercial use — “For commercial production rights when shipping generated audio, see Commercial Licensing.” source
- API — “Looking to integrate the Bulbul V3 API within your Products/Applications?” source
Security
- SOC 2 Type II — “ISO 27001 and SOC 2 Type II certified.” source
- SOC 2 — “ISO 27001 and SOC 2 Type II certified.” source
- ISO 27001 — “We're ISO 27001:2022 certified and hold SOC 2 Type II.” source
- Not: ISO 42001 — “ISO 42001 is in progress, and we'll publish it when it's done.” source
- Not: Not trained on your data — “Custom models trained on a customer's data remain inside the customer's environment, with weights they own and we never reuse.” source
Latest updates
- Sarvam Vision 2.1: Pushing the Pareto frontier of document intelligence (2.1)
Sarvam Vision 2.1 adds structured extraction and Indic handwritten recognition capabilities.
- Introducing Saaras V4 (V4)
Introduces Saaras V4, an ASR model for a multilingual world.
- Everything we announced at Sarvam Epoch
Announces new models, Sarvam Inference, Indus agents, Sarvam Code, and more.
Coqui XTTS
Coqui XTTS is a multilingual text-to-speech model that clones a voice from a short audio clip and generates speech.
Plans
PRO
- $9 /month
- 10× private storage capacity
- 2× public storage capacity
- 20× included inference credits
- 8× ZeroGPU quota and highest queue priority
- Host ZeroGPU, Gradio & Docker Spaces
- Spaces Dev Mode
Team
- $20 /month per user
- SSO support (SAML & OIDC)
- Data location control with Storage Regions
- Detailed action reviews with Audit Logs
- Granular access control via Resource Groups
- Repository usage Analytics
- Advanced auth policies and repository visibility controls
Enterprise
- $50 /month per user
- + All benefits from the Team plan
- Highest storage, bandwidth, and API rate limits
- Automated user management with SCIM provisioning
- Advanced security and access controls
- Managed billing with annual commitments
- Dedicated support
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Text to speech — “Text-to-Speech” source
- Voice cloning — “Voice cloning with just a 6-second audio clip.” source
- Many languages — “Supports 17 languages.” source
- API — “Using 🐸TTS API:” source