tools
Fish Audio Review 2026 — Open-Source Voice Cloning and Real-Time TTS
Fish Audio delivers ElevenLabs-quality voice cloning for free. Open-source Fish Speech, real-time TTS API, Docker deployment, 8 languages. Best self-hosted voice AI 2026.

In inglese
Fish Audio is a web platform and API for generating speech from text, cloning voices, designing voices from prompts, converting speech to text, changing voices, translating audio, and creating longer-form audio projects. It is used by creators, developers, publishers, educators, game teams, and voice-agent builders.
The free plan has monthly credit and generation limits. API access uses separate pay-as-you-go pricing, while commercial rights and production guarantees vary by plan and use case.
Features
- Generate speech from text with expressive voice controls
- Clone voices from reference audio
- Design custom voices from text prompts
- Convert speech to text with an API
- Stream TTS and ASR over WebSocket
- Create voiceovers, audiobooks, and longer audio projects
- Translate audio and separate voices from other sounds
- Open-weight Fish models require a paid commercial license
Use cases
- Create voiceovers for videos, ads, and explainers
- Build real-time voice agents and conversational applications
- Produce audiobooks and long-form narration
- Generate game and character dialogue
- Localize content while preserving a speaker's voice
- Prototype speech features with the Fish Audio API
Pros
Cons
Latest updates
- Fish Audio S2 (Fish Audio S2)
Launched the S2 text-to-speech model with inline emotion cues, multi-speaker dialogue, 80+ languages, and API model ID s2-pro.
- Fish Audio S1 (Fish Audio S1)
Rebranded Fish Speech to Fish Audio and introduced S1 and S1-mini TTS models with multilingual support and emotional expressions.
- v1.5.1 (v1.5.1)
Added ONNX export and Arabic and Hebrew text processing; fixed PyTorch security settings and Apple Silicon compatibility.
- v1.5.0 (v1.5.0)
Introduced the v1.5 model architecture, bearer-token API authentication, reference-audio caching by hash, and base64 reference data in JSON.
- v1.4.3 (v1.4.3)
Introduced Fish Agent for conversational AI with streaming and real-time interactions, and fixed non-English speech issues.
Capabilities
- Text to speech — “Text to Speech with the most natural and human sounding AI voice generator” source
- Voice cloning — “Instant voice cloning from 10 seconds of audio, 60+ emotion tags, sub-300ms streaming latency, and 500,000+ community voices.” source
- Speech to text — “Speech to Text” source
- Sound effects — “Sound Effects” source
- Dubbing — “Audio Translation” source
- Many languages — “Speak 30+ languages with any voice” source
- Commercial use — “Premium subscribers can use verified voices (that you own) for commercial purposes.” source
- API — “Yes! Premium subscribers get access to our flexible pay-as-you-go API.” source
Get it
Pricing
- Starting price
- $5.50/mo
- Prices checked
- 2026-09-24
Free Tier
Free
- 8,000 credits monthly
- Up to 7 minutes generation
- Up to 500 characters per generation
- 3 public voice slots
- Standard generation speed
- Enhanced voice cloning
Plus
- $5.5 mo
- $66 billed annually
- 15 mo
- 250,000 credits monthly
- Up to 200 minutes generation
- Up to 15,000 characters per generation
- Unlimited public + 10 private voice slots
- Priority generation on our latest models
- Access to Voice Design
Pro
- $37.5 mo
- $450 billed annually
- 2,000,000 credits monthly
- Up to 1,620 minutes generation
- 3 team seats included
- Up to 30,000 characters per generation
- Unlimited voice slots
- 5 professional voice slots
Max
- $749 mo
- $8988 billed annually
- 25,000,000 credits monthly
- Up to 6,250 minutes generation
- 10 team seats included
- 15 professional voice slots
- And everything in Pro
Enterprise
Price on request
- Pay as you go with organization-level controls
- Zero Data Retention
- On-Premise Deployment
- SOC2 Compliance
- Custom SSO (Coming soon)