Tools
Fish Audio vs Rime Arcana
Fish Audio or Rime Arcana? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In inglese
In short
- Both: Text to speech, Voice cloning, Many languages, API
- Only Fish Audio states: Speech to text, Sound effects, Dubbing, Commercial use
Fish Audio
Fish Audio turns text into expressive speech, with voice cloning, voice design, speech-to-text, audio tools, and developer APIs.
Plans
Free Tier
Free
- 8,000 credits monthly
- Up to 7 minutes generation
- Up to 500 characters per generation
- 3 public voice slots
- Standard generation speed
- Enhanced voice cloning
Plus
- $5.5 mo
- $66 billed annually
- 15 mo
- 250,000 credits monthly
- Up to 200 minutes generation
- Up to 15,000 characters per generation
- Unlimited public + 10 private voice slots
- Priority generation on our latest models
- Access to Voice Design
Pro
- $37.5 mo
- $450 billed annually
- 2,000,000 credits monthly
- Up to 1,620 minutes generation
- 3 team seats included
- Up to 30,000 characters per generation
- Unlimited voice slots
- 5 professional voice slots
Max
- $749 mo
- $8988 billed annually
- 25,000,000 credits monthly
- Up to 6,250 minutes generation
- 10 team seats included
- 15 professional voice slots
- And everything in Pro
Enterprise
Price on request
- Pay as you go with organization-level controls
- Zero Data Retention
- On-Premise Deployment
- SOC2 Compliance
- Custom SSO (Coming soon)
Prices checked 2026-09-24 on the maker’s page.
Capabilities
- Text to speech — “Text to Speech with the most natural and human sounding AI voice generator” source
- Voice cloning — “Instant voice cloning from 10 seconds of audio, 60+ emotion tags, sub-300ms streaming latency, and 500,000+ community voices.” source
- Speech to text — “Speech to Text” source
- Sound effects — “Sound Effects” source
- Dubbing — “Audio Translation” source
- Many languages — “Speak 30+ languages with any voice” source
- Commercial use — “Premium subscribers can use verified voices (that you own) for commercial purposes.” source
- API — “Yes! Premium subscribers get access to our flexible pay-as-you-go API.” source
Latest updates
- Fish Audio S2 (Fish Audio S2)
Launched the S2 text-to-speech model with inline emotion cues, multi-speaker dialogue, 80+ languages, and API model ID s2-pro.
- Fish Audio S1 (Fish Audio S1)
Rebranded Fish Speech to Fish Audio and introduced S1 and S1-mini TTS models with multilingual support and emotional expressions.
- v1.5.1 (v1.5.1)
Added ONNX export and Arabic and Hebrew text processing; fixed PyTorch security settings and Apple Silicon compatibility.
Rime Arcana
Rime Arcana is a text-to-speech model for real-time voice apps, generating expressive multilingual speech through Rime’s API.
Plans
Starter
- starting at $0.03 / 1K characters
- $0.05 / 1K characters
- ~800 minutes free (about 800k characters)
- 20 concurrent TTS generations
- Public Slack support
Enterprise
Price on request
- Unlimited concurrent TTS generations
- Unlimited custom TTS voice clones
- SLAs + dedicated support
- Cloud, on-prem, or VPC
- BAA (HIPAA) and SOC 2 reports
Prices checked 2026-09-26 on the maker’s page.
Capabilities
- Text to speech — “Arcana is a multimodal, autoregressive text-to-speech (TTS) model that generates discrete audio tokens from text inputs.” source
- Voice cloning — “Unlimited custom TTS voice clones” source
- Many languages — “Arcana v3's TTS AI supports 10 languages out of the box, with more coming soon.” source
- API — “It’s production-ready and API-accessible from day one.” source
Security
- SOC 2 Type II — “As of May 2025, we are SOC 2 Type 2 compliant .” source
- SOC 2 — “As of May 2025, we are SOC 2 Type 2 compliant .” source
- HIPAA — “As of February 2024, we are HIPAA compliant .” source
- Not trained on your data — “We never use customer text or audio to train our models” source
Latest updates
- Arcana is Now Unlimited
Removal of character limits enables longer speech passages; Arcana accommodates WAV, PCM, and MULAW, adjustable sampling rates, and text normalization.