Instructions to use prasadvittaldev/svanita-0.6b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prasadvittaldev/svanita-0.6b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="prasadvittaldev/svanita-0.6b")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("prasadvittaldev/svanita-0.6b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Svanita — speech-to-text for Indian languages, on a CPU
Svanita is a speech-to-text model family for South and Southeast Asian languages, starting with India. It is built for how people actually speak: mixing English words into their own language, over app audio and phone lines, with the transcript produced fast enough on an ordinary CPU for a live voice agent.
This release (v0.2) covers Telugu and Hindi in one checkpoint. More Indian languages follow on the same recipe.
If you only need Telugu, use the
v0.1branch. That Telugu-only checkpoint is still 1–3 WER points better on every Telugu test set than this multilingual one. The honest numbers are in Evaluation.
Svanita transcribes the local language and the English mixed into it in one transcript — native script for the native words, Latin script for the English ones:
OK home loan కావాలి అంటే మీకు ఎంత కావాలి
ये तो आम का best रहता है और एक mix जो आता है mix vegetable वाला वो best रहता है
It runs faster than real time on an ordinary desktop CPU, with no GPU: a 3.6-second reply is transcribed in about 0.33 seconds.
svanita_v02_telugu_hindi_demo.mp4 — six clips from speakers it never heard in training, three Telugu and three Hindi, transcribed live on a CPU (including the two it gets wrong), then why the architecture is cheap on a CPU. The older Telugu-only demo is svanita_demo.mp4.
Built by Prasad Vittaldev. Follow on LinkedIn for new languages and voices.
Why Svanita exists
Conversation across India is code-switched. In Telugu: "documents submit చేయాల్సి ఉంటుంది bank వచ్చి మీరు". In Hindi: "आप मिलिएगा sir". Existing open recognizers write those English words in the native script (డాక్యుమెంట్స్ సబ్మిట్, सर). That is accurate as phonetics, but a downstream LLM, search index or form-filler reads it as an unknown native word. Svanita writes them as English, which is what a voice agent actually needs.
The second goal was CPU real time. Svanita is built on NVIDIA's Parakeet TDT, whose architecture is unusually cheap to run on a CPU (see How it is fast on a CPU).
Model
| Architecture | FastConformer encoder (24 layers) + TDT (token-and-duration transducer) decoder |
| Parameters | ~622M (encoder 609M, decoder + joint ~13M) |
| Languages in this release | Telugu and Hindi, with code-switched English |
| Base model | nvidia/parakeet-tdt-0.6b-v3 |
| Vocabulary | new 4,097-piece tokenizer covering Telugu script, Devanagari and Latin (the base model's vocabulary has neither Indian script) |
| Input | 16 kHz mono audio (wideband or telephone quality) |
| Output | native script, with English words in Latin script; lowercase except acronyms (OK, OTP); no punctuation |
| Licence | CC-BY-4.0 |
Usage
Tested with transformers 5.12.1 and torch 2.11.
Fast CPU transcription (recommended)
transcribe_cpu.py is a single-file greedy TDT decoder over the model's own modules — the loop the speed figures below were measured with. It also takes --lang, which locks the output to one Indian script:
pip install transformers torch soundfile
python transcribe_cpu.py sample_1.wav --lang telugu --threads 8
# OK home loan కావాలి అంటే మీకు ఎంత కావాలి
# [3.6 s of audio transcribed in 326 ms on CPU, 8 threads]
python transcribe_cpu.py sample_hi_2.wav --lang hindi --threads 8
# ये तो आम का best रहता है और एक mix जो आता है mix vegetable वाला वो best रहता है
# [5.1 s of audio transcribed in 451 ms on CPU, 8 threads]
Pass --lang whenever you know the language. One multilingual checkpoint sometimes answers Telugu audio in Devanagari, and Hindi audio in Telugu script — 8% and 4% of test clips respectively. --lang removes every token of the other Indian script from the vocabulary before decoding (English stays in Latin), which is worth 2.0 WER points on Telugu and 0.6 on Hindi, at no cost in speed. A voice agent always knows the language of the call.
Five sample clips are included — Telugu (sample_1.wav, sample_3.wav, sample_5.wav) and Hindi (sample_hi_1.wav, sample_hi_2.wav) — all from held-out test speakers.
With transformers
import torch, soundfile as sf
from transformers import ParakeetForTDT, ParakeetProcessor
repo = "prasadvittaldev/svanita-0.6b"
processor = ParakeetProcessor.from_pretrained(repo)
model = ParakeetForTDT.from_pretrained(repo).eval()
audio, sr = sf.read("sample_3.wav", dtype="float32") # 16 kHz mono
inputs = processor(audio, sampling_rate=16000)
with torch.inference_mode():
out = model.generate(**inputs)
print(processor.batch_decode(out.sequences, skip_special_tokens=True)[0])
# మళ్ళీ దానికి కూడా documents submit చేయాల్సి ఉంటుంది bank వచ్చి మీరు
generate() has no script lock; use transcribe_cpu.py for that.
Evaluation
All scores are on speakers never seen in training, with --lang set. The same normalization is applied to every system (Unicode NFC, zero-width characters removed, punctuation stripped, Latin lowercased).
Two WER figures are reported, because they answer different questions:
- WER (either script) counts an English word as correct whether the system wrote it in Latin (
order) or the native script (ఆర్డర్). This is the fairest measure of recognition across systems. - WER (Latin required) counts it as correct only in Latin. This measures what Svanita is designed for: text a downstream system can read as English.
Hindi
IndicVoices Hindi, conversational / extempore speech — 2,599 clips, 4.19 h, 248 speakers:
| system | WER (either script) | WER (Latin required) | CER | English kept in Latin |
|---|---|---|---|---|
| Svanita v0.2 | 23.7 | 23.8 | 13.6 | 73% |
| IndicConformer 600M, RNN-T | 15.5 | 20.6 | 16.1 | 0% |
| IndicConformer 600M, CTC | 17.1 | 22.0 | 16.4 | 0% |
| Whisper large-v2, Hindi fine-tune | 37.3 | 40.0 | 33.2 | 0% |
| slice | Svanita v0.2 WER / CER | IndicConformer RNN-T | Whisper-Hindi large-v2 |
|---|---|---|---|
| IndicVoices, 8 kHz telephone copy | 26.5 / 16.0 | 17.6 / 17.5 | 45.2 / 41.2 |
| IndicVoices, code-mixed clips only (1,136) | 24.0 / 14.4 (Latin required: 24.1) | 16.2 / 24.3 (Latin required: 25.7) | 39.3 / 39.9 (Latin required: 44.2) |
| Kathbath, read speech (3,151 clips, 20 speakers) | 25.4 / 12.6 | 8.9 / 2.8 | 11.7 / 4.3 |
| Kathbath, 8 kHz telephone copy | 28.7 / 15.1 | 11.8 / 3.9 | — |
Telugu
IndicVoices Telugu, conversational / extempore speech — 1,499 clips, 2.34 h, 153 speakers:
| system | WER (either script) | WER (Latin required) | CER | English kept in Latin |
|---|---|---|---|---|
| Svanita v0.2 (this checkpoint) | 37.7 | 38.0 | 17.0 | 74% |
Svanita v0.1 (Telugu-only, v0.1 branch) |
36.5 | 36.9 | 16.4 | 76% |
| IndicConformer 600M, RNN-T | 27.4 | 39.1 | 22.1 | 0% |
| IndicConformer 600M, CTC | 29.0 | 40.3 | 22.1 | 0% |
| Whisper large-v2, Telugu fine-tune | 44.0 | 52.3 | 32.0 | 0% |
| slice | Svanita v0.2 WER / CER | Svanita v0.1 | IndicConformer RNN-T |
|---|---|---|---|
| IndicVoices, 8 kHz telephone copy | 40.7 / 19.2 | 38.8 / 18.2 | 31.1 / 23.5 |
| IndicVoices, code-mixed clips only (800) | 38.8 / 19.8 | 37.6 / 19.2 | 29.8 / 33.5 |
| Kathbath, read speech (2,378 clips, 20 speakers) | 39.7 / 10.1 | 36.3 / 8.9 | 20.9 / 3.1 |
| Kathbath, 8 kHz telephone copy | 43.1 / 12.4 | 40.2 / 11.3 | 23.6 / 4.2 |
Baselines: ai4bharat/indic-conformer-600m-multilingual, vasista22/whisper-telugu-large-v2 and vasista22/whisper-hindi-large-v2, run on the same audio with the same scoring.
Reading the results honestly:
- IndicConformer recognizes more words correctly in both languages, clearly so on read speech, and is the stronger recognizer overall. Svanita is ahead on character accuracy in Hindi, on keeping English as English (it wins every Latin-required comparison), and against the Whisper fine-tunes on conversational and telephone speech by 13–19 points.
- Adding Hindi cost Telugu 1–3 points. Hindi is 60% of the training audio, and the shared vocabulary gives Telugu fewer word pieces than the Telugu-only v0.1 had. A Telugu-weighted top-up run recovered about half of the loss; the rest remains. For Telugu-only deployments,
v0.1is the better checkpoint. - Short turns are the hardest case for every system.
How it is fast on a CPU
- It only processes the audio it is given. The encoder downsamples 8×, so one second of speech is just 13 frames and a 3-second reply is 38. Whisper pads every input to a fixed 30-second window (1,500 frames) regardless of length.
- The big network runs once. The 609M-parameter encoder reads the whole clip in a single pass; only the ~13M-parameter decoder and joint network loop while the text is written. About 98% of the model runs once per utterance.
- It jumps over frames. A token-and-duration transducer predicts, at each step, a word-piece and how many frames to skip (0–4), so pauses and long sounds are crossed in one hop instead of frame by frame.
Measured latency — AMD Ryzen 5 8600G (6 cores / 12 threads), fp32, no GPU, transcribe_cpu.py, 8 threads:
| turn length | time |
|---|---|
| 3.6 s | 326 ms |
| 3.9 s | 387 ms |
| 5.1 s | 451 ms |
| 6.4 s | 515 ms |
That is 10–12× faster than real time. For comparison, Whisper large-v2 fine-tunes reached about 4× real time on a GPU in our evaluation.
Keep it in full precision. In our tests, PyTorch dynamic int8 quantization made it about twice as fast but raised WER by ~30 points, and an ONNX export was fast but not numerically faithful. The fp32 model above is what we recommend and what the numbers describe.
Training
Data — 1,136 hours, 625,184 clips:
| language | hours | clips | speakers |
|---|---|---|---|
| Telugu | 451 | 241,291 | 1,691 |
| Hindi | 685 | 383,893 | 1,941 |
ai4bharat/IndicVoices(Telugu and Hindi): conversational, extempore and read speech from many districts. Its annotations mark English words spoken inside the native sentence (टाइम [time]); Svanita's targets use the Latin form, so 46% of Telugu and 39% of Hindi training clips contain Latin-script English. Transcriber annotations ([noise],[uhh],[stammers], …) are removed.ai4bharat/Kathbath(Telugu and Hindi): read speech.- Held-out evaluation speakers were removed from training entirely (125,297 Hindi and 76,692 Telugu training clips from speakers who also appear in the validation splits were dropped).
- About 40% of training clips were degraded on the fly to telephone quality (300–3,400 Hz, 8 kHz, μ-law/8-bit quantization, 15–30 dB noise), so one model serves app audio and phone calls. SpecAugment was also applied.
Recipe
- Vocabulary swap. The base model's tokenizer has neither Indian script. A 4,097-piece Unigram tokenizer was trained on the Telugu + Hindi + Latin transcripts, and only the two vocabulary-shaped tensors were re-initialized (decoder embedding and joint output head, ~5.3M parameters). The pretrained duration outputs, the 609M encoder and the decoder LSTM were kept — the encoder warm-started from the Telugu-only v0.1 checkpoint.
- Stage A — encoder frozen, decoder and joint trained: 27,500 steps, learning rate 1e-3, 8-bit AdamW, best checkpoint at 17,500.
- Stage B — full fine-tune from Stage A's best checkpoint: 61,815 steps (2 epochs), learning rate 3e-4 with the encoder at 3e-5, warmup 2,000, cosine decay, bf16 autocast, gradient checkpointing, gradient clipping at 1.0. Best checkpoint at 52,500.
- Telugu-weighted top-up — 12,000 further steps on a 70% Telugu / 30% Hindi mix at learning rate 1e-4, to win back accuracy Telugu lost to the bigger Hindi share. This checkpoint is step 10,000 of that run, selected by Telugu dev WER; it recovered about half the loss without moving Hindi.
Trained on a single NVIDIA RTX 5060 Ti (16 GB).
Limitations
- Accuracy is short of the best open recognizers for both languages (IndicConformer). It is suitable for intent-level understanding by a downstream LLM, not for text a person will read verbatim.
- Telugu is better served by
v0.1if you do not need Hindi. - Read speech is a weakness (Kathbath: 25.4% Hindi / 39.7% Telugu, against IndicConformer's 8.9% / 20.9%). Most of the training audio is conversational; only 87 hours of Hindi and 81 hours of Telugu are read speech. More read speech is the clearest next improvement.
- Pass
--lang. Without it, the model occasionally answers in the other Indian script (8% of Telugu clips, 4% of Hindi). - Pure English sentences are not supported well. English was learned only as words inside Indian-language speech; on all-English clips WER is high and some English words, especially Indian names, come out in the native script.
- English words fused to native suffixes (
డేస్లో, "days-lo") are usually written in the native script, because the training annotations do not split them.
- Downloads last month
- 333
Model tree for prasadvittaldev/svanita-0.6b
Base model
nvidia/parakeet-tdt-0.6b-v3Datasets used to train prasadvittaldev/svanita-0.6b
ai4bharat/Kathbath
Evaluation results
- WER (either script accepted) on IndicVoices Hindi (held-out speakers)self-reported23.700
- WER (English must be in Latin script) on IndicVoices Hindi (held-out speakers)self-reported23.800
- CER on IndicVoices Hindi (held-out speakers)self-reported13.600
- WER (either script accepted) on IndicVoices Telugu (held-out speakers)self-reported37.700
- WER (English must be in Latin script) on IndicVoices Telugu (held-out speakers)self-reported38.000
- CER on IndicVoices Telugu (held-out speakers)self-reported17.000
- WER on Kathbath Hindi (read speech)self-reported25.400
- CER on Kathbath Hindi (read speech)self-reported12.600