CoolFace
Modelpublic

RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian

sourceHugging Faceopenmdw-1.1updated 29d agoView on Hugging Face
4likes395downloads
Model Card

Nemotron 3.5 ASR — Jordanian Dialect Arabic

Production-ready streaming ASR for Jordanian dialect Arabic, built for voice agents. A full fine-tune of nvidia/nemotron-3.5-asr-streaming-0.6b on ~10 hours of Jordanian speech.

WER 39.48% → 28.82% — a 27% relative gain, and 33% on CER — on a held-out 3.32-hour test set, at RTF 0.107 with 560 ms chunks on a single L40S.

python
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition",
                model="RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian")
print(pipe("audio.wav"))

Why this model

The accuracy gain is free. RTF is identical to the base model at every chunk size — 0.107 vs 0.107 at 560 ms. Weight values changed; shapes, parameter count, dtype and decode path did not. A 27% relative WER improvement costs nothing in serving.

Three failure modes eliminated, not just reduced. The base model appended a spurious language tag to a third of its outputs and emitted wrong-script text on another third. Both go to zero. Empty hypotheses drop from 100 to 8. A pipeline on the base model needs a tag stripper, a script filter and an empty-output fallback; this one needs none of them.

One checkpoint, a 14× latency range. 80 ms to 1120 ms is a runtime argument, not a retrain. Shift the latency/throughput operating point per request without holding a second model in memory.

Zero-config Arabic. default_prompt_id is pinned to Arabic, so the one-liner above is the complete API. No language argument, no manifest, no prompt key to get wrong.

Two runtimes, one repo. Ships in both 🤗 Transformers and NeMo format from the same verified weights.

Cheap to reproduce. 66.9 minutes on a single L40S — about 1.1 GPU-hours — for the full 10-epoch run.


Model details

ArchitectureFastConformer-CacheAware-RNNT with language-ID prompt conditioning
EncoderCache-aware FastConformer, 24 layers, d\_model 1024, 8× conv subsampling (one frame = 80 ms)
DecoderRNN-T, prediction net 640, 13,087-piece SentencePiece vocab
Parameters637M (all updated during fine-tuning)
LanguageArabic — Jordanian dialect (ar, prompt index 7, pinned as default_prompt_id)
Sampling rate16 kHz mono
Checkpointnemotron35-jordanian.nemo (2.55 GB) · model.safetensors (2.55 GB)
Chunk sizes80 / 160 / 320 / 560 / 1120 ms, selectable at runtime

Native punctuation and capitalisation. The RNN-T decoder is monotonic, so output is emitted as audio arrives rather than after the utterance completes.

Both formats carry identical weights — the safetensors file was converted from the .nemo and verified elementwise, then re-scored on the same test set — so every number below applies to either path.


Accuracy

Test set: 2,067 utterances, 3.32 hours, never trained or validated on. Base and fine-tuned scored with the same helper, same normaliser, same att_context_size = [56, 3], same batch size, same GPU. The only variable is the checkpoint.

Headline — normalised, language tags stripped:

BaseFine-tunedAbsoluteRelative
WER39.48%28.82%−10.66−27.0%
CER17.00%11.39%−5.61−33.0%

Drop-in unmodified — raw, as decoded:

BaseFine-tunedAbsolute
WER47.17%30.10%−17.07
CER22.21%11.81%−10.40

The gain holds under every scoring convention:

ConventionBase WERFT WERBase CERFT CER
Raw, as decoded47.17%30.10%22.21%11.81%
Language tags stripped45.18%30.10%18.73%11.81%
Normalised, tags kept43.89%28.82%19.52%11.39%
Normalised + tags stripped39.48%28.82%17.00%11.39%

Normalisation folds Arabic orthographic variation (diacritics, alef/ya/ta-marbuta forms) and punctuation, applied identically to reference and hypothesis.

Output quality

The bigger practical win is not the WER.

BaseFine-tuned
Empty hypotheses1008
Utterances with a spurious <xx-XX> language tag687 (33.2%)0
Utterances containing non-Arabic script720 (34.8%)0

Every language tag scored as an insertion against references that contain none — roughly 4.4 WER points of pure artefact, now gone. The remaining error is real transcription error, which makes the number meaningful rather than inflated.

Numbers are written as words

Output spells numerals out in Arabic rather than emitting digits — تسعين, not ٩٠ or 90. This is consistent, so it is straightforward to handle, but it is not what most ASR output looks like:

SpokenOutputNot
ninetyتسعين٩٠ / 90
twenty-fiveخمسة وعشرين٢٥ / 25

Applications that need numeric values — phone numbers, order IDs, quantities, prices — should run an Arabic word-to-number pass over the transcript. Scoring is unaffected as long as references follow the same convention.

NVIDIA reports ar-AR at 12.55% WER on FLEURS at 320 ms. The gap to 39–47% for the base here is dialect and channel, not misconfiguration — the base was run with the Arabic prompt, not auto. FLEURS is read MSA; this test set is spontaneous Jordanian. That gap is precisely what the fine-tune closes.

Speed

Single stream, cache-aware per-chunk inference, L40S, 100 utterances / 13.7 minutes:

`att_context_size`ChunkBase WERFT WERFT CERBase RTFFT RTFCompute/chunkEnd-to-end latency
[56, 0]80 ms47.27%36.08%15.08%0.6050.59847.6 ms~128 ms
[56, 1]\*160 ms46.01%34.51%14.77%0.3290.31950.5 ms~211 ms
[56, 3]320 ms51.37%37.18%15.89%0.1700.17053.2 ms~373 ms
`[56, 6]`560 ms46.80%34.77%14.51%0.1070.10757.6 ms~618 ms
[56, 13]1120 ms47.69%33.67%13.42%0.0620.06365.2 ms~1185 ms

\* Not in the checkpoint's declared context list — validate before deploying.

At the recommended 560 ms setting, one stream occupies roughly 11% of an L40S. Offline batch reaches RTF 0.0031 at batch 8. Both are single-model, single-GPU measurements on the hardware stated.

RTF numbers only compare across identical hardware, batch size, precision and chunk size.


Usage

Audio must be 16 kHz mono.

🤗 Transformers

bash
pip install "transformers>=5.13.0"
python
from transformers import pipeline

pipe = pipeline("automatic-speech-recognition",
                model="RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian")
print(pipe("audio.wav"))

default_prompt_id is pinned to Arabic (index 7) rather than the base model's auto-detect (101), so no language argument is needed and the spurious <xx-XX> tags cannot return.

Defaults to 560 ms chunks. To change the latency/accuracy operating point:

python
from transformers import AutoProcessor, AutoModelForRNNT
from transformers.audio_utils import load_audio

model_id  = "RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian"
processor = AutoProcessor.from_pretrained(model_id)
model     = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")

processor.set_num_lookahead_tokens(6)   # 0/3/6/13 = 80/320/560/1120 ms

sr     = processor.feature_extractor.sampling_rate
audio  = load_audio("audio.wav", sampling_rate=sr)
inputs = processor(audio, sampling_rate=sr).to(model.device, dtype=model.dtype)
out    = model.generate(**inputs, return_dict_in_generate=True)
print(processor.decode(out.sequences, skip_special_tokens=True))

processor.decode returns a list of strings, one per utterance in the batch.

NeMo

bash
apt-get update && apt-get install -y libsndfile1 ffmpeg
pip install "nemo_toolkit[asr]==3.0.0" "numba-cuda[cu12]" "numpy==2.2.6" \
    jiwer soundfile huggingface_hub

Installation issues: see the NVIDIA guide.

The NeMo path is prompt-conditioned through the manifest — it reads the language from the lang field, so a bare file path raises ValueError: Unknown prompt key: 'None'. This does not apply to the Transformers path above, where the language is baked into the config.

python
import json, tempfile, os, torch
from huggingface_hub import hf_hub_download
from nemo.collections.asr.models import ASRModel

CKPT = hf_hub_download(
    "RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian",
    "nemotron35-jordanian.nemo",
)

model = ASRModel.restore_from(CKPT, map_location="cuda")
model.eval()
model.encoder.set_default_att_context_size([56, 6])  # 560 ms chunks

def transcribe(model, wav_paths, batch_size=8):
    rows = [{"audio_filepath": p, "duration": 0.0, "text": "", "lang": "ar"}
            for p in wav_paths]
    with tempfile.NamedTemporaryFile("w", suffix=".json", delete=False,
                                     encoding="utf-8") as tmp:
        for r in rows:
            tmp.write(json.dumps(r, ensure_ascii=False) + "\n")
        path = tmp.name
    try:
        with torch.inference_mode():
            out = model.transcribe(path, batch_size=batch_size)
    finally:
        os.unlink(path)
    if isinstance(out, tuple):
        out = out[0]
    return [h.text if hasattr(h, "text") else str(h) for h in out]

print(transcribe(model, ["sample.wav"], batch_size=1)[0])

Evaluate on a manifest

One JSON object per line with audio_filepath, duration, text, lang. duration is in seconds; lang must be "ar" on every row.

json
{"audio_filepath": "/data/test/utt_0001.wav", "duration": 3.42, "text": "أهلا وسهلا كيف بقدر أساعدك", "lang": "ar"}
{"audio_filepath": "/data/test/utt_0002.wav", "duration": 5.18, "text": "بدي أستفسر عن الطلبية اللي عملتها مبارح", "lang": "ar"}
{"audio_filepath": "/data/test/utt_0003.wav", "duration": 1.07, "text": "تفضل", "lang": "ar"}
python
import json, time
from nemo.collections.asr.metrics.wer import word_error_rate

rows = [json.loads(l) for l in open("test_manifest.json", encoding="utf-8") if l.strip()]
refs = [r["text"] for r in rows]

t0 = time.time()
hyps = transcribe(model, [r["audio_filepath"] for r in rows], batch_size=8)
elapsed = time.time() - t0
audio_s = sum(r["duration"] for r in rows)

print(f"WER {word_error_rate(hypotheses=hyps, references=refs)*100:.2f}%")
print(f"CER {word_error_rate(hypotheses=hyps, references=refs, use_cer=True)*100:.2f}%")
print(f"RTF {elapsed/audio_s:.4f}")

For Arabic, apply the same normalisation to references and hypotheses before scoring — folding diacritics, alef/ya/ta-marbuta variants and punctuation — or the number is not comparable to anyone else's. Keep the normaliser's Arabic character classes written as \uXXXX escapes: as literal Arabic, bidirectional reordering in an editor silently swaps the endpoints of a character range, and the normaliser then deletes the text it was meant to fold.

Streaming

bash
python examples/asr/asr_cache_aware_streaming/speech_to_text_cache_aware_streaming_infer.py \
    model_path=/path/to/nemotron35-jordanian.nemo \
    dataset_manifest=test_manifest.json \
    batch_size=1 \
    target_lang=ar \
    att_context_size="[56,6]" \
    strip_lang_tags=true \
    output_path=./stream_out

Second value of att_context_size sets latency: 0 = 80 ms, 1 = 160 ms, 3 = 320 ms, 6 = 560 ms, 13 = 1120 ms. Left context stays at 56 frames.

Hardcode ar. Do not use target_lang=auto — it reintroduces the tag emission this fine-tune eliminated. Keep strip_lang_tags=true as insurance on out-of-distribution audio.


Scope and intended use

Built for voice agents: intent classification, slot filling, keyword spotting — tasks where a downstream language model absorbs word-level errors and the 27% relative gain translates directly into better intent accuracy.

Numeric slots need a conversion pass. Numbers come out as Arabic words, not digits, so phone numbers, order IDs and amounts require word-to-number post-processing. See Numbers are written as words.

Not for verbatim transcription. At 28.8% WER, output should not be read unedited or used for compliance recording.

Code-switching is the hardest category at 36.45% WER (n=99), though also where fine-tuning helped most — the base scored 56.87%. Single borrowed words transliterate into Arabic (ال manager → المانيجر), which is arguably correct but scores as a substitution. Continuous English degrades further. Worth knowing if your calls open with a scripted English greeting.

Corpus-level gain, not universal. 226 of 1,311 utterances regressed after fine-tuning; the headline number is a large improvement on most, net of real regression on some.

Domain scope. Tuned for Jordanian dialect on this corpus's channel conditions and transcription conventions. The test set comes from a different sector than the training data, with different speakers and topics — so in-domain performance is likely better than these numbers suggest. Performance on MSA or other Arabic dialects is unmeasured.


Training

Method: full fine-tune, all 637M parameters updated, none frozen. NeMo has no first-class LoRA path for FastConformer-RNNT, and the model fits on one L40S at this batch size. Tokenizer kept from the base checkpoint.

SplitUtterancesHours
Train6,89510.35
Validation1610.47
Test2,0673.32

Validation carved out of train and split by recording ID — a per-utterance split would leak speaker and session across the boundary. Test never trained or validated on.

Hyperparameters:

OptimizerAdamW, betas [0.9, 0.98], weight decay 1e-3
SchedulerNoamAnnealing, warmup 500 steps, min\_lr 1e-6
optim.lr0.1 → effective peak LR 1.40e-4
Gradient clipping0.5
Precisionbf16
BatchingLhotse dynamic bucketing, batch\duration 100 s, quadratic\duration 15, 10 buckets
Epochs10 (9,200 steps)
SpecAugment2 freq masks (width 27), 10 time masks (width 0.05)
Prompt modelangID, always the real language id

Hardware: single NVIDIA L40S. Python 3.12.6, torch 2.8.0+cu129, NeMo 3.0.0, numpy 2.2.6. 66.9 minutes for 10 epochs ≈ 1.1 GPU-hours.

Fine-tuning this base model yourself — three config traps

  1. 1.optim.lr is a Noam multiplier, not a learning rate: peak = lr × 1.40e-3. The default 2.0 gives 6.3e-4, too hot for a warm checkpoint; a sane-looking 1e-4 gives 1.4e-7 and trains on nothing.
  2. 2.optim.sched.d_model must be a literal 1024. The YAML interpolation is detached by the time Lightning calls configure_optimizers → InterpolationKeyError.
  3. 3.use_bucketing defaults to false, so num_buckets is silently ignored. is_tarred must be false for loose .wav files.

Training data

Trained on a private Jordanian dialect Arabic corpus, not publicly available. The test set is private, so the numbers here are reported for base-vs-fine-tuned comparison under identical conditions rather than as an independently reproducible benchmark.


License

Governed by OpenMDW-1.1, inherited from nvidia/nemotron-3.5-asr-streaming-0.6b.


Citation

bibtex
@misc{nemotron35_jordanian,
  title  = {Nemotron 3.5 ASR fine-tuned for Jordanian Dialect Arabic},
  author = {FILL IN},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian}}
}

Base model:

bibtex
@misc{nemotron35asr,
  title  = {Nemotron 3.5 ASR},
  author = {NVIDIA},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b}}
}

References: Stateful Conformer with Cache-based Inference · Fast Conformer


Maintainer: FILL IN · Questions and issues via the Community tab.