Armenian FastConformer (ONNX)
This is an ONNX export of nvidia/stt_hy_fastconformer_hybrid_large_pc, NVIDIA's Armenian speech recognition model with punctuation and capitalization. It exports the transducer (RNNT) branch. You can run it with onnx-asr or any runtime that loads NeMo Parakeet ONNX models.
It was made for Tetro, a local meeting transcriber for macOS.
Files
| File | Size | Notes |
|---|---|---|
encoder-model.int8.onnx |
125 MB | Dynamic int8 quantization (onnxruntime quantize_dynamic) |
decoder_joint-model.int8.onnx |
5 MB | Dynamic int8 quantization |
encoder-model.onnx |
435 MB | Full precision |
decoder_joint-model.onnx |
20 MB | Full precision |
nemo80.onnx |
87 KB | 80-band log-mel preprocessor, taken from onnx-asr |
vocab.txt |
13 KB | 1024 SentencePiece tokens plus <blk> (index 1024) |
config.json |
onnx-asr config: nemo-conformer-rnnt, 80 features, subsampling 8 |
Usage
import onnx_asr
model = onnx_asr.load_model("nemo-conformer-rnnt", "path/to/this/folder", quantization="int8")
print(model.recognize("armenian.wav")) # 16 kHz mono
Checks
On three Common Voice hy-AM test clips, NeMo's own RNNT decoder and this ONNX export gave the same text at both full precision and int8. That confirms the export matches NeMo. It says nothing about accuracy, which is inherited from the base model; see its model card for results. The model sometimes writes "եւ" where you might expect "և".
How it was made
The model was exported with NeMo (set_export_config({"decoder_type": "rnnt"}), then export), and the int8 files were produced with onnxruntime dynamic quantization. The scripts are in scripts/.
License
This model is CC-BY-4.0, the same license as the base model. Credit: NVIDIA, stt_hy_fastconformer_hybrid_large_pc.
- Downloads last month
- 85
Model tree for Vana-Labs/stt-fastconformer-armenian
Base model
nvidia/stt_hy_fastconformer_hybrid_large_pc