tools
Saaras
Saaras is Sarvam AI's speech-to-text model for transcribing, translating, and processing speech in Indian languages.

Saaras converts speech to text through REST, WebSocket, and batch APIs. It supports transcription, translation, verbatim output, transliteration, and code-mixed speech across Indian languages and English.
It is designed for developers building voice assistants, call-center tools, captioning systems, and multilingual applications. Usage is billed by audio duration; batch diarization costs more. It is not presented as a standalone consumer chat app.
Features
- Speech-to-text for 22 Indian languages and English
- Transcribe, translate, verbatim, translit, and codemix output modes
- REST, WebSocket, and batch processing endpoints
- Streaming speech recognition
- Batch transcription for longer audio files
- Batch diarization with speaker identification
- Supports code-mixed and regional speech
Use cases
- Build multilingual voice assistants
- Transcribe customer-support and call-center recordings
- Generate captions for Indian-language audio and video
- Translate spoken content between supported languages
- Analyze speakers in recorded conversations
Pros
Cons
Pricing
- Starting price
- ₹30.00 per hour
- Pricing checked
- 2026-09-19
Real-time
₹30.00 per hour
- Speech-to-text API
Streaming
₹30.00 per hour
- Streaming speech-to-text
Batch
₹30.00 per hour
- Batch speech-to-text
Batch with diarization
₹45.00 per hour
- Batch speech-to-text
- Speaker identification