Armenian FastConformer (ONNX)

This is an ONNX export of nvidia/stt_hy_fastconformer_hybrid_large_pc, NVIDIA's Armenian speech recognition model with punctuation and capitalization. It exports the transducer (RNNT) branch. You can run it with onnx-asr or any runtime that loads NeMo Parakeet ONNX models.

It was made for Tetro, a local meeting transcriber for macOS.

Files

File Size Notes
encoder-model.int8.onnx 125 MB Dynamic int8 quantization (onnxruntime quantize_dynamic)
decoder_joint-model.int8.onnx 5 MB Dynamic int8 quantization
encoder-model.onnx 435 MB Full precision
decoder_joint-model.onnx 20 MB Full precision
nemo80.onnx 87 KB 80-band log-mel preprocessor, taken from onnx-asr
vocab.txt 13 KB 1024 SentencePiece tokens plus <blk> (index 1024)
config.json onnx-asr config: nemo-conformer-rnnt, 80 features, subsampling 8

Usage

import onnx_asr
model = onnx_asr.load_model("nemo-conformer-rnnt", "path/to/this/folder", quantization="int8")
print(model.recognize("armenian.wav"))  # 16 kHz mono

Checks

On three Common Voice hy-AM test clips, NeMo's own RNNT decoder and this ONNX export gave the same text at both full precision and int8. That confirms the export matches NeMo. It says nothing about accuracy, which is inherited from the base model; see its model card for results. The model sometimes writes "եւ" where you might expect "և".

How it was made

The model was exported with NeMo (set_export_config({"decoder_type": "rnnt"}), then export), and the int8 files were produced with onnxruntime dynamic quantization. The scripts are in scripts/.

License

This model is CC-BY-4.0, the same license as the base model. Credit: NVIDIA, stt_hy_fastconformer_hybrid_large_pc.

Downloads last month
85
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vana-Labs/stt-fastconformer-armenian

Quantized
(3)
this model