Proposal: Native Sindhi (سنڌي) and Roman Sindhi Support for Jev-Urdu
Hi Muhammad Noman (@muhammadnoman76 ),
First of all, congratulations on building Jev-Urdu! The decision-oriented, System One architecture on top of ModernBERT/mmBERT is a brilliant, lightweight alternative to bulky generative LLMs for practical routing and triage workflows.
I am an NLP researcher and developer working on low-resource Pakistani languages (specifically Sindhi and Roman Sindhi). Over the past few days, I conducted an empirical study exploring how well the mmBERT base and the Jev decision head generalize to Perso-Arabic Sindhi and Roman Sindhi.
What I Discovered (Zero-Shot Baseline):
- Multilingual Tokenizer: Because of mmBERT’s 256K vocabulary, Sindhi script and Roman Sindhi phonetics are already represented as clean subwords with almost zero
<unk>fallback. - Sindhi Script Decisions: The base model zero-shot scored 87.5% on Sindhi sentiment, 75.0% on triage urgency, and 71.9% on topic classification.
- Roman Sindhi & Cross-Script Alignment: Roman Sindhi achieved 66.7% zero-shot when matched with Latin options (
positive,negative,neutral), with 100% precision on cognate roots.
What I Implemented in my Fork (shakeel143/jev-multilingual):
- Native Sindhi Task Presets (
SINDHI_PRESETS): Added natural Sindhi presets forsentiment,triage,check_claim, andtopicinsidelibrary/src/jev_urdu/api.py. - Language-Aware API (100% Backwards-Compatible):
- Callers can use
model.sentiment(text, lang="sd"),model.triage(text, lang="sd"), andmodel.topic(text, lang="sd"). - Existing Urdu calls default to
lang="ur"and remain completely unaffected.
- Callers can use
- Comprehensive Unit Tests: Extended
library/tests/test_runtime.pyandconftest.pywith full offline test fixture support (all 22 unit tests passing). - Low-Resource Decision-Head Adaptation: Fine-tuned the decision head with token embeddings strictly frozen (
train_embeddings: false), which successfully resolved the zero-shot antonym contradiction bug (گرم ↔ ٿڌي) while shielding Urdu from regression (0% forgetting).
Proposed Contribution:
I would love to contribute these Sindhi & Roman Sindhi presets and test suites upstream into jev-urdu so that the library natively serves both major Pakistani national languages (Urdu and Sindhi).
- Fork Repository: shakeel143/jev-multilingual
- Code Changes: Minimal 4-file diff adding
SINDHI_PRESETS,langkwargs, and unit tests without changing core inference mechanics.
Would you be open to a Pull Request merging these native Sindhi presets and tests into jev-urdu? I'd be honored to collaborate on expanding the multilingual footprint of LughaatNLP!
Best regards,
Shakeel Ahmed Sanjrani
Hugging Face Profile: @shakeel143