Proposal: Native Sindhi (سنڌي) and Roman Sindhi Support for Jev-Urdu

#1
by shakeel143 - opened

Hi Muhammad Noman (@muhammadnoman76 ),

First of all, congratulations on building Jev-Urdu! The decision-oriented, System One architecture on top of ModernBERT/mmBERT is a brilliant, lightweight alternative to bulky generative LLMs for practical routing and triage workflows.

I am an NLP researcher and developer working on low-resource Pakistani languages (specifically Sindhi and Roman Sindhi). Over the past few days, I conducted an empirical study exploring how well the mmBERT base and the Jev decision head generalize to Perso-Arabic Sindhi and Roman Sindhi.

What I Discovered (Zero-Shot Baseline):

  1. Multilingual Tokenizer: Because of mmBERT’s 256K vocabulary, Sindhi script and Roman Sindhi phonetics are already represented as clean subwords with almost zero <unk> fallback.
  2. Sindhi Script Decisions: The base model zero-shot scored 87.5% on Sindhi sentiment, 75.0% on triage urgency, and 71.9% on topic classification.
  3. Roman Sindhi & Cross-Script Alignment: Roman Sindhi achieved 66.7% zero-shot when matched with Latin options (positive, negative, neutral), with 100% precision on cognate roots.

What I Implemented in my Fork (shakeel143/jev-multilingual):

  1. Native Sindhi Task Presets (SINDHI_PRESETS): Added natural Sindhi presets for sentiment, triage, check_claim, and topic inside library/src/jev_urdu/api.py.
  2. Language-Aware API (100% Backwards-Compatible):
    • Callers can use model.sentiment(text, lang="sd"), model.triage(text, lang="sd"), and model.topic(text, lang="sd").
    • Existing Urdu calls default to lang="ur" and remain completely unaffected.
  3. Comprehensive Unit Tests: Extended library/tests/test_runtime.py and conftest.py with full offline test fixture support (all 22 unit tests passing).
  4. Low-Resource Decision-Head Adaptation: Fine-tuned the decision head with token embeddings strictly frozen (train_embeddings: false), which successfully resolved the zero-shot antonym contradiction bug (گرم ↔ ٿڌي) while shielding Urdu from regression (0% forgetting).

Proposed Contribution:

I would love to contribute these Sindhi & Roman Sindhi presets and test suites upstream into jev-urdu so that the library natively serves both major Pakistani national languages (Urdu and Sindhi).

  • Fork Repository: shakeel143/jev-multilingual
  • Code Changes: Minimal 4-file diff adding SINDHI_PRESETS, lang kwargs, and unit tests without changing core inference mechanics.

Would you be open to a Pull Request merging these native Sindhi presets and tests into jev-urdu? I'd be honored to collaborate on expanding the multilingual footprint of LughaatNLP!

Best regards,
Shakeel Ahmed Sanjrani
Hugging Face Profile: @shakeel143

Sign up or log in to comment