Tools
Coqui XTTS vs Rime Arcana
Coqui XTTS or Rime Arcana? Their plans monthly and yearly, the capabilities their makers state, security and latest updates, side by side — read from the makers' own pages.
In short
- Both: Text to speech, Voice cloning, Many languages, API
Coqui XTTS
Coqui XTTS is a multilingual text-to-speech model that clones a voice from a short audio clip and generates speech.
Plans
PRO
- $9 /month
- 10× private storage capacity
- 2× public storage capacity
- 20× included inference credits
- 8× ZeroGPU quota and highest queue priority
- Host ZeroGPU, Gradio & Docker Spaces
- Spaces Dev Mode
Team
- $20 /month per user
- SSO support (SAML & OIDC)
- Data location control with Storage Regions
- Detailed action reviews with Audit Logs
- Granular access control via Resource Groups
- Repository usage Analytics
- Advanced auth policies and repository visibility controls
Enterprise
- $50 /month per user
- + All benefits from the Team plan
- Highest storage, bandwidth, and API rate limits
- Automated user management with SCIM provisioning
- Advanced security and access controls
- Managed billing with annual commitments
- Dedicated support
Prices checked 2026-09-25 on the maker’s page.
Capabilities
- Text to speech — “Text-to-Speech” source
- Voice cloning — “Voice cloning with just a 6-second audio clip.” source
- Many languages — “Supports 17 languages.” source
- API — “Using 🐸TTS API:” source
Rime Arcana
Rime Arcana is a text-to-speech model for real-time voice apps, generating expressive multilingual speech through Rime’s API.
Plans
Starter
- starting at $0.03 / 1K characters
- $0.05 / 1K characters
- ~800 minutes free (about 800k characters)
- 20 concurrent TTS generations
- Public Slack support
Enterprise
Price on request
- Unlimited concurrent TTS generations
- Unlimited custom TTS voice clones
- SLAs + dedicated support
- Cloud, on-prem, or VPC
- BAA (HIPAA) and SOC 2 reports
Prices checked 2026-09-26 on the maker’s page.
Capabilities
- Text to speech — “Arcana is a multimodal, autoregressive text-to-speech (TTS) model that generates discrete audio tokens from text inputs.” source
- Voice cloning — “Unlimited custom TTS voice clones” source
- Many languages — “Arcana v3's TTS AI supports 10 languages out of the box, with more coming soon.” source
- API — “It’s production-ready and API-accessible from day one.” source
Security
- SOC 2 Type II — “As of May 2025, we are SOC 2 Type 2 compliant .” source
- SOC 2 — “As of May 2025, we are SOC 2 Type 2 compliant .” source
- HIPAA — “As of February 2024, we are HIPAA compliant .” source
- Not trained on your data — “We never use customer text or audio to train our models” source
Latest updates
- Arcana is Now Unlimited
Removal of character limits enables longer speech passages; Arcana accommodates WAV, PCM, and MULAW, adjustable sampling rates, and text normalization.