Tools
Coqui XTTS vs Fish Audio
Coqui XTTS or Fish Audio? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In inglese
In short
- Both: Text to speech, Voice cloning, Many languages, API
- Only Fish Audio states: Speech to text, Sound effects, Dubbing, Commercial use
Coqui XTTS
Coqui XTTS is a multilingual text-to-speech model that clones a voice from a short audio clip and generates speech.
Plans
PRO
- $9 /month
- 10× private storage capacity
- 2× public storage capacity
- 20× included inference credits
- 8× ZeroGPU quota and highest queue priority
- Host ZeroGPU, Gradio & Docker Spaces
- Spaces Dev Mode
Team
- $20 /month per user
- SSO support (SAML & OIDC)
- Data location control with Storage Regions
- Detailed action reviews with Audit Logs
- Granular access control via Resource Groups
- Repository usage Analytics
- Advanced auth policies and repository visibility controls
Enterprise
- $50 /month per user
- + All benefits from the Team plan
- Highest storage, bandwidth, and API rate limits
- Automated user management with SCIM provisioning
- Advanced security and access controls
- Managed billing with annual commitments
- Dedicated support
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Text to speech — “Text-to-Speech” source
- Voice cloning — “Voice cloning with just a 6-second audio clip.” source
- Many languages — “Supports 17 languages.” source
- API — “Using 🐸TTS API:” source
Fish Audio
Fish Audio turns text into expressive speech, with voice cloning, voice design, speech-to-text, audio tools, and developer APIs.
Plans
Free Tier
Free
- 8,000 credits monthly
- Up to 7 minutes generation
- Up to 500 characters per generation
- 3 public voice slots
- Standard generation speed
- Enhanced voice cloning
Plus
- $5.5 mo
- $66 billed annually
- 15 mo
- 250,000 credits monthly
- Up to 200 minutes generation
- Up to 15,000 characters per generation
- Unlimited public + 10 private voice slots
- Priority generation on our latest models
- Access to Voice Design
Pro
- $37.5 mo
- $450 billed annually
- 2,000,000 credits monthly
- Up to 1,620 minutes generation
- 3 team seats included
- Up to 30,000 characters per generation
- Unlimited voice slots
- 5 professional voice slots
Max
- $749 mo
- $8988 billed annually
- 25,000,000 credits monthly
- Up to 6,250 minutes generation
- 10 team seats included
- 15 professional voice slots
- And everything in Pro
Enterprise
Price on request
- Pay as you go with organization-level controls
- Zero Data Retention
- On-Premise Deployment
- SOC2 Compliance
- Custom SSO (Coming soon)
Prices checked 2026-09-24 on the maker’s page.
Capabilities
- Text to speech — “Text to Speech with the most natural and human sounding AI voice generator” source
- Voice cloning — “Instant voice cloning from 10 seconds of audio, 60+ emotion tags, sub-300ms streaming latency, and 500,000+ community voices.” source
- Speech to text — “Speech to Text” source
- Sound effects — “Sound Effects” source
- Dubbing — “Audio Translation” source
- Many languages — “Speak 30+ languages with any voice” source
- Commercial use — “Premium subscribers can use verified voices (that you own) for commercial purposes.” source
- API — “Yes! Premium subscribers get access to our flexible pay-as-you-go API.” source
Latest updates
- Fish Audio S2 (Fish Audio S2)
Launched the S2 text-to-speech model with inline emotion cues, multi-speaker dialogue, 80+ languages, and API model ID s2-pro.
- Fish Audio S1 (Fish Audio S1)
Rebranded Fish Speech to Fish Audio and introduced S1 and S1-mini TTS models with multilingual support and emotional expressions.
- v1.5.1 (v1.5.1)
Added ONNX export and Arabic and Hebrew text processing; fixed PyTorch security settings and Apple Silicon compatibility.