Skip to content
AI.info

tools

Saaras

Saaras is Sarvam AI's speech-to-text model for transcribing, translating, and processing speech in Indian languages.

Saaras

Saaras converts speech to text through REST, WebSocket, and batch APIs. It supports transcription, translation, verbatim output, transliteration, and code-mixed speech across Indian languages and English.

It is designed for developers building voice assistants, call-center tools, captioning systems, and multilingual applications. Usage is billed by audio duration; batch diarization costs more. It is not presented as a standalone consumer chat app.

Features

  • Speech-to-text for 22 Indian languages and English
  • Transcribe, translate, verbatim, translit, and codemix output modes
  • REST, WebSocket, and batch processing endpoints
  • Streaming speech recognition
  • Batch transcription for longer audio files
  • Batch diarization with speaker identification
  • Supports code-mixed and regional speech

Use cases

  • Build multilingual voice assistants
  • Transcribe customer-support and call-center recordings
  • Generate captions for Indian-language audio and video
  • Translate spoken content between supported languages
  • Analyze speakers in recorded conversations

Pros

    Cons

      Pricing

      Starting price
      ₹30.00 per hour
      Pricing checked
      2026-09-19

      Real-time

      ₹30.00 per hour

      • Speech-to-text API

      Streaming

      ₹30.00 per hour

      • Streaming speech-to-text

      Batch

      ₹30.00 per hour

      • Batch speech-to-text

      Batch with diarization

      ₹45.00 per hour

      • Batch speech-to-text
      • Speaker identification
      Official website