Svanita — speech-to-text for Indian languages, on a CPU

Svanita is a speech-to-text model family for South and Southeast Asian languages, starting with India. It is built for how people actually speak: mixing English words into their own language, over app audio and phone lines, with the transcript produced fast enough on an ordinary CPU for a live voice agent.

This release (v0.2) covers Telugu and Hindi in one checkpoint. More Indian languages follow on the same recipe.

If you only need Telugu, use the v0.1 branch. That Telugu-only checkpoint is still 1–3 WER points better on every Telugu test set than this multilingual one. The honest numbers are in Evaluation.

Svanita transcribes the local language and the English mixed into it in one transcript — native script for the native words, Latin script for the English ones:

OK home loan కావాలి అంటే మీకు ఎంత కావాలి

ये तो आम का best रहता है और एक mix जो आता है mix vegetable वाला वो best रहता है

It runs faster than real time on an ordinary desktop CPU, with no GPU: a 3.6-second reply is transcribed in about 0.33 seconds.

svanita_v02_telugu_hindi_demo.mp4 — six clips from speakers it never heard in training, three Telugu and three Hindi, transcribed live on a CPU (including the two it gets wrong), then why the architecture is cheap on a CPU. The older Telugu-only demo is svanita_demo.mp4.

Built by Prasad Vittaldev. Follow on LinkedIn for new languages and voices.

Why Svanita exists

Conversation across India is code-switched. In Telugu: "documents submit చేయాల్సి ఉంటుంది bank వచ్చి మీరు". In Hindi: "आप मिलिएगा sir". Existing open recognizers write those English words in the native script (డాక్యుమెంట్స్ సబ్మిట్, सर). That is accurate as phonetics, but a downstream LLM, search index or form-filler reads it as an unknown native word. Svanita writes them as English, which is what a voice agent actually needs.

The second goal was CPU real time. Svanita is built on NVIDIA's Parakeet TDT, whose architecture is unusually cheap to run on a CPU (see How it is fast on a CPU).

Model

Architecture FastConformer encoder (24 layers) + TDT (token-and-duration transducer) decoder
Parameters ~622M (encoder 609M, decoder + joint ~13M)
Languages in this release Telugu and Hindi, with code-switched English
Base model nvidia/parakeet-tdt-0.6b-v3
Vocabulary new 4,097-piece tokenizer covering Telugu script, Devanagari and Latin (the base model's vocabulary has neither Indian script)
Input 16 kHz mono audio (wideband or telephone quality)
Output native script, with English words in Latin script; lowercase except acronyms (OK, OTP); no punctuation
Licence CC-BY-4.0

Usage

Tested with transformers 5.12.1 and torch 2.11.

Fast CPU transcription (recommended)

transcribe_cpu.py is a single-file greedy TDT decoder over the model's own modules — the loop the speed figures below were measured with. It also takes --lang, which locks the output to one Indian script:

pip install transformers torch soundfile
python transcribe_cpu.py sample_1.wav --lang telugu --threads 8
# OK home loan కావాలి అంటే మీకు ఎంత కావాలి
# [3.6 s of audio transcribed in 326 ms on CPU, 8 threads]

python transcribe_cpu.py sample_hi_2.wav --lang hindi --threads 8
# ये तो आम का best रहता है और एक mix जो आता है mix vegetable वाला वो best रहता है
# [5.1 s of audio transcribed in 451 ms on CPU, 8 threads]

Pass --lang whenever you know the language. One multilingual checkpoint sometimes answers Telugu audio in Devanagari, and Hindi audio in Telugu script — 8% and 4% of test clips respectively. --lang removes every token of the other Indian script from the vocabulary before decoding (English stays in Latin), which is worth 2.0 WER points on Telugu and 0.6 on Hindi, at no cost in speed. A voice agent always knows the language of the call.

Five sample clips are included — Telugu (sample_1.wav, sample_3.wav, sample_5.wav) and Hindi (sample_hi_1.wav, sample_hi_2.wav) — all from held-out test speakers.

With transformers

import torch, soundfile as sf
from transformers import ParakeetForTDT, ParakeetProcessor

repo = "prasadvittaldev/svanita-0.6b"
processor = ParakeetProcessor.from_pretrained(repo)
model = ParakeetForTDT.from_pretrained(repo).eval()

audio, sr = sf.read("sample_3.wav", dtype="float32")        # 16 kHz mono
inputs = processor(audio, sampling_rate=16000)
with torch.inference_mode():
    out = model.generate(**inputs)
print(processor.batch_decode(out.sequences, skip_special_tokens=True)[0])
# మళ్ళీ దానికి కూడా documents submit చేయాల్సి ఉంటుంది bank వచ్చి మీరు

generate() has no script lock; use transcribe_cpu.py for that.

Evaluation

All scores are on speakers never seen in training, with --lang set. The same normalization is applied to every system (Unicode NFC, zero-width characters removed, punctuation stripped, Latin lowercased).

Two WER figures are reported, because they answer different questions:

  • WER (either script) counts an English word as correct whether the system wrote it in Latin (order) or the native script (ఆర్డర్). This is the fairest measure of recognition across systems.
  • WER (Latin required) counts it as correct only in Latin. This measures what Svanita is designed for: text a downstream system can read as English.

Hindi

IndicVoices Hindi, conversational / extempore speech — 2,599 clips, 4.19 h, 248 speakers:

system WER (either script) WER (Latin required) CER English kept in Latin
Svanita v0.2 23.7 23.8 13.6 73%
IndicConformer 600M, RNN-T 15.5 20.6 16.1 0%
IndicConformer 600M, CTC 17.1 22.0 16.4 0%
Whisper large-v2, Hindi fine-tune 37.3 40.0 33.2 0%
slice Svanita v0.2 WER / CER IndicConformer RNN-T Whisper-Hindi large-v2
IndicVoices, 8 kHz telephone copy 26.5 / 16.0 17.6 / 17.5 45.2 / 41.2
IndicVoices, code-mixed clips only (1,136) 24.0 / 14.4 (Latin required: 24.1) 16.2 / 24.3 (Latin required: 25.7) 39.3 / 39.9 (Latin required: 44.2)
Kathbath, read speech (3,151 clips, 20 speakers) 25.4 / 12.6 8.9 / 2.8 11.7 / 4.3
Kathbath, 8 kHz telephone copy 28.7 / 15.1 11.8 / 3.9 —

Telugu

IndicVoices Telugu, conversational / extempore speech — 1,499 clips, 2.34 h, 153 speakers:

system WER (either script) WER (Latin required) CER English kept in Latin
Svanita v0.2 (this checkpoint) 37.7 38.0 17.0 74%
Svanita v0.1 (Telugu-only, v0.1 branch) 36.5 36.9 16.4 76%
IndicConformer 600M, RNN-T 27.4 39.1 22.1 0%
IndicConformer 600M, CTC 29.0 40.3 22.1 0%
Whisper large-v2, Telugu fine-tune 44.0 52.3 32.0 0%
slice Svanita v0.2 WER / CER Svanita v0.1 IndicConformer RNN-T
IndicVoices, 8 kHz telephone copy 40.7 / 19.2 38.8 / 18.2 31.1 / 23.5
IndicVoices, code-mixed clips only (800) 38.8 / 19.8 37.6 / 19.2 29.8 / 33.5
Kathbath, read speech (2,378 clips, 20 speakers) 39.7 / 10.1 36.3 / 8.9 20.9 / 3.1
Kathbath, 8 kHz telephone copy 43.1 / 12.4 40.2 / 11.3 23.6 / 4.2

Baselines: ai4bharat/indic-conformer-600m-multilingual, vasista22/whisper-telugu-large-v2 and vasista22/whisper-hindi-large-v2, run on the same audio with the same scoring.

Reading the results honestly:

  • IndicConformer recognizes more words correctly in both languages, clearly so on read speech, and is the stronger recognizer overall. Svanita is ahead on character accuracy in Hindi, on keeping English as English (it wins every Latin-required comparison), and against the Whisper fine-tunes on conversational and telephone speech by 13–19 points.
  • Adding Hindi cost Telugu 1–3 points. Hindi is 60% of the training audio, and the shared vocabulary gives Telugu fewer word pieces than the Telugu-only v0.1 had. A Telugu-weighted top-up run recovered about half of the loss; the rest remains. For Telugu-only deployments, v0.1 is the better checkpoint.
  • Short turns are the hardest case for every system.

How it is fast on a CPU

  1. It only processes the audio it is given. The encoder downsamples 8×, so one second of speech is just 13 frames and a 3-second reply is 38. Whisper pads every input to a fixed 30-second window (1,500 frames) regardless of length.
  2. The big network runs once. The 609M-parameter encoder reads the whole clip in a single pass; only the ~13M-parameter decoder and joint network loop while the text is written. About 98% of the model runs once per utterance.
  3. It jumps over frames. A token-and-duration transducer predicts, at each step, a word-piece and how many frames to skip (0–4), so pauses and long sounds are crossed in one hop instead of frame by frame.

Measured latency — AMD Ryzen 5 8600G (6 cores / 12 threads), fp32, no GPU, transcribe_cpu.py, 8 threads:

turn length time
3.6 s 326 ms
3.9 s 387 ms
5.1 s 451 ms
6.4 s 515 ms

That is 10–12× faster than real time. For comparison, Whisper large-v2 fine-tunes reached about 4× real time on a GPU in our evaluation.

Keep it in full precision. In our tests, PyTorch dynamic int8 quantization made it about twice as fast but raised WER by ~30 points, and an ONNX export was fast but not numerically faithful. The fp32 model above is what we recommend and what the numbers describe.

Training

Data — 1,136 hours, 625,184 clips:

language hours clips speakers
Telugu 451 241,291 1,691
Hindi 685 383,893 1,941
  • ai4bharat/IndicVoices (Telugu and Hindi): conversational, extempore and read speech from many districts. Its annotations mark English words spoken inside the native sentence (टाइम [time]); Svanita's targets use the Latin form, so 46% of Telugu and 39% of Hindi training clips contain Latin-script English. Transcriber annotations ([noise], [uhh], [stammers], …) are removed.
  • ai4bharat/Kathbath (Telugu and Hindi): read speech.
  • Held-out evaluation speakers were removed from training entirely (125,297 Hindi and 76,692 Telugu training clips from speakers who also appear in the validation splits were dropped).
  • About 40% of training clips were degraded on the fly to telephone quality (300–3,400 Hz, 8 kHz, μ-law/8-bit quantization, 15–30 dB noise), so one model serves app audio and phone calls. SpecAugment was also applied.

Recipe

  1. Vocabulary swap. The base model's tokenizer has neither Indian script. A 4,097-piece Unigram tokenizer was trained on the Telugu + Hindi + Latin transcripts, and only the two vocabulary-shaped tensors were re-initialized (decoder embedding and joint output head, ~5.3M parameters). The pretrained duration outputs, the 609M encoder and the decoder LSTM were kept — the encoder warm-started from the Telugu-only v0.1 checkpoint.
  2. Stage A — encoder frozen, decoder and joint trained: 27,500 steps, learning rate 1e-3, 8-bit AdamW, best checkpoint at 17,500.
  3. Stage B — full fine-tune from Stage A's best checkpoint: 61,815 steps (2 epochs), learning rate 3e-4 with the encoder at 3e-5, warmup 2,000, cosine decay, bf16 autocast, gradient checkpointing, gradient clipping at 1.0. Best checkpoint at 52,500.
  4. Telugu-weighted top-up — 12,000 further steps on a 70% Telugu / 30% Hindi mix at learning rate 1e-4, to win back accuracy Telugu lost to the bigger Hindi share. This checkpoint is step 10,000 of that run, selected by Telugu dev WER; it recovered about half the loss without moving Hindi.

Trained on a single NVIDIA RTX 5060 Ti (16 GB).

Limitations

  • Accuracy is short of the best open recognizers for both languages (IndicConformer). It is suitable for intent-level understanding by a downstream LLM, not for text a person will read verbatim.
  • Telugu is better served by v0.1 if you do not need Hindi.
  • Read speech is a weakness (Kathbath: 25.4% Hindi / 39.7% Telugu, against IndicConformer's 8.9% / 20.9%). Most of the training audio is conversational; only 87 hours of Hindi and 81 hours of Telugu are read speech. More read speech is the clearest next improvement.
  • Pass --lang. Without it, the model occasionally answers in the other Indian script (8% of Telugu clips, 4% of Hindi).
  • Pure English sentences are not supported well. English was learned only as words inside Indian-language speech; on all-English clips WER is high and some English words, especially Indian names, come out in the native script.
  • English words fused to native suffixes (డేస్‍లో, "days-lo") are usually written in the native script, because the training annotations do not split them.
Downloads last month
333
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prasadvittaldev/svanita-0.6b

Finetuned
(107)
this model

Datasets used to train prasadvittaldev/svanita-0.6b

Evaluation results

  • WER (either script accepted) on IndicVoices Hindi (held-out speakers)
    self-reported
    23.700
  • WER (English must be in Latin script) on IndicVoices Hindi (held-out speakers)
    self-reported
    23.800
  • CER on IndicVoices Hindi (held-out speakers)
    self-reported
    13.600
  • WER (either script accepted) on IndicVoices Telugu (held-out speakers)
    self-reported
    37.700
  • WER (English must be in Latin script) on IndicVoices Telugu (held-out speakers)
    self-reported
    38.000
  • CER on IndicVoices Telugu (held-out speakers)
    self-reported
    17.000
  • WER on Kathbath Hindi (read speech)
    self-reported
    25.400
  • CER on Kathbath Hindi (read speech)
    self-reported
    12.600