Skip to content
AI.info

Tools

Coqui XTTS vs Fish Audio

Coqui XTTS or Fish Audio? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.

In short

  • Both: Text to speech, Voice cloning, Many languages, API
  • Only Fish Audio states: Speech to text, Sound effects, Dubbing, Commercial use

Coqui XTTS

Coqui XTTS is a multilingual text-to-speech model that clones a voice from a short audio clip and generates speech.

Plans

PRO

  • $9 /month
  • 10× private storage capacity
  • 2× public storage capacity
  • 20× included inference credits
  • 8× ZeroGPU quota and highest queue priority
  • Host ZeroGPU, Gradio & Docker Spaces
  • Spaces Dev Mode

Team

  • $20 /month per user
  • SSO support (SAML & OIDC)
  • Data location control with Storage Regions
  • Detailed action reviews with Audit Logs
  • Granular access control via Resource Groups
  • Repository usage Analytics
  • Advanced auth policies and repository visibility controls

Enterprise

  • $50 /month per user
  • + All benefits from the Team plan
  • Highest storage, bandwidth, and API rate limits
  • Automated user management with SCIM provisioning
  • Advanced security and access controls
  • Managed billing with annual commitments
  • Dedicated support

Prices checked 2026-09-25 on the maker’s page.

Capabilities

  • Text to speech — “Text-to-Speech” source
  • Voice cloning — “Voice cloning with just a 6-second audio clip.” source
  • Many languages — “Supports 17 languages.” source
  • API — “Using 🐸TTS API:” source
About Coqui XTTS

Fish Audio

Fish Audio turns text into expressive speech, with voice cloning, voice design, speech-to-text, audio tools, and developer APIs.

Plans

Free Tier

Free

  • 8,000 credits monthly
  • Up to 7 minutes generation
  • Up to 500 characters per generation
  • 3 public voice slots
  • Standard generation speed
  • Enhanced voice cloning

Plus

  • $5.5 mo
  • $66 billed annually
  • 15 mo
  • 250,000 credits monthly
  • Up to 200 minutes generation
  • Up to 15,000 characters per generation
  • Unlimited public + 10 private voice slots
  • Priority generation on our latest models
  • Access to Voice Design

Pro

  • $37.5 mo
  • $450 billed annually
  • 2,000,000 credits monthly
  • Up to 1,620 minutes generation
  • 3 team seats included
  • Up to 30,000 characters per generation
  • Unlimited voice slots
  • 5 professional voice slots

Max

  • $749 mo
  • $8988 billed annually
  • 25,000,000 credits monthly
  • Up to 6,250 minutes generation
  • 10 team seats included
  • 15 professional voice slots
  • And everything in Pro

Enterprise

Price on request

  • Pay as you go with organization-level controls
  • Zero Data Retention
  • On-Premise Deployment
  • SOC2 Compliance
  • Custom SSO (Coming soon)

Prices checked 2026-09-24 on the maker’s page.

Capabilities

  • Text to speech — “Text to Speech with the most natural and human sounding AI voice generator” source
  • Voice cloning — “Instant voice cloning from 10 seconds of audio, 60+ emotion tags, sub-300ms streaming latency, and 500,000+ community voices.” source
  • Speech to text — “Speech to Text” source
  • Sound effects — “Sound Effects” source
  • Dubbing — “Audio Translation” source
  • Many languages — “Speak 30+ languages with any voice” source
  • Commercial use — “Premium subscribers can use verified voices (that you own) for commercial purposes.” source
  • API — “Yes! Premium subscribers get access to our flexible pay-as-you-go API.” source

Latest updates

  • Fish Audio S2 (Fish Audio S2)

    Launched the S2 text-to-speech model with inline emotion cues, multi-speaker dialogue, 80+ languages, and API model ID s2-pro.

  • Fish Audio S1 (Fish Audio S1)

    Rebranded Fish Speech to Fish Audio and introduced S1 and S1-mini TTS models with multilingual support and emotional expressions.

  • v1.5.1 (v1.5.1)

    Added ONNX export and Arabic and Hebrew text processing; fixed PyTorch security settings and Apple Silicon compatibility.

About Fish Audio