RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian
Nemotron 3.5 ASR — Jordanian Dialect Arabic
Production-ready streaming ASR for Jordanian dialect Arabic, built for voice agents. A full fine-tune of nvidia/nemotron-3.5-asr-streaming-0.6b on ~10 hours of Jordanian speech.
WER 39.48% → 28.82% — a 27% relative gain, and 33% on CER — on a held-out 3.32-hour test set, at RTF 0.107 with 560 ms chunks on a single L40S.
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition",
model="RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian")
print(pipe("audio.wav"))Why this model
The accuracy gain is free. RTF is identical to the base model at every chunk size — 0.107 vs 0.107 at 560 ms. Weight values changed; shapes, parameter count, dtype and decode path did not. A 27% relative WER improvement costs nothing in serving.
Three failure modes eliminated, not just reduced. The base model appended a spurious language tag to a third of its outputs and emitted wrong-script text on another third. Both go to zero. Empty hypotheses drop from 100 to 8. A pipeline on the base model needs a tag stripper, a script filter and an empty-output fallback; this one needs none of them.
One checkpoint, a 14× latency range. 80 ms to 1120 ms is a runtime argument, not a retrain. Shift the latency/throughput operating point per request without holding a second model in memory.
Zero-config Arabic. default_prompt_id is pinned to Arabic, so the one-liner above is the complete API. No language argument, no manifest, no prompt key to get wrong.
Two runtimes, one repo. Ships in both 🤗 Transformers and NeMo format from the same verified weights.
Cheap to reproduce. 66.9 minutes on a single L40S — about 1.1 GPU-hours — for the full 10-epoch run.
Model details
Native punctuation and capitalisation. The RNN-T decoder is monotonic, so output is emitted as audio arrives rather than after the utterance completes.
Both formats carry identical weights — the safetensors file was converted from the .nemo and verified elementwise, then re-scored on the same test set — so every number below applies to either path.
Accuracy
Test set: 2,067 utterances, 3.32 hours, never trained or validated on. Base and fine-tuned scored with the same helper, same normaliser, same att_context_size = [56, 3], same batch size, same GPU. The only variable is the checkpoint.
Headline — normalised, language tags stripped:
Drop-in unmodified — raw, as decoded:
The gain holds under every scoring convention:
Normalisation folds Arabic orthographic variation (diacritics, alef/ya/ta-marbuta forms) and punctuation, applied identically to reference and hypothesis.
Output quality
The bigger practical win is not the WER.
Every language tag scored as an insertion against references that contain none — roughly 4.4 WER points of pure artefact, now gone. The remaining error is real transcription error, which makes the number meaningful rather than inflated.
Numbers are written as words
Output spells numerals out in Arabic rather than emitting digits — تسعين, not ٩٠ or 90. This is consistent, so it is straightforward to handle, but it is not what most ASR output looks like:
Applications that need numeric values — phone numbers, order IDs, quantities, prices — should run an Arabic word-to-number pass over the transcript. Scoring is unaffected as long as references follow the same convention.
NVIDIA reports ar-AR at 12.55% WER on FLEURS at 320 ms. The gap to 39–47% for the base here is dialect and channel, not misconfiguration — the base was run with the Arabic prompt, not auto. FLEURS is read MSA; this test set is spontaneous Jordanian. That gap is precisely what the fine-tune closes.Speed
Single stream, cache-aware per-chunk inference, L40S, 100 utterances / 13.7 minutes:
\* Not in the checkpoint's declared context list — validate before deploying.
At the recommended 560 ms setting, one stream occupies roughly 11% of an L40S. Offline batch reaches RTF 0.0031 at batch 8. Both are single-model, single-GPU measurements on the hardware stated.
RTF numbers only compare across identical hardware, batch size, precision and chunk size.
Usage
Audio must be 16 kHz mono.
🤗 Transformers
pip install "transformers>=5.13.0"from transformers import pipeline
pipe = pipeline("automatic-speech-recognition",
model="RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian")
print(pipe("audio.wav"))default_prompt_id is pinned to Arabic (index 7) rather than the base model's auto-detect (101), so no language argument is needed and the spurious <xx-XX> tags cannot return.
Defaults to 560 ms chunks. To change the latency/accuracy operating point:
from transformers import AutoProcessor, AutoModelForRNNT
from transformers.audio_utils import load_audio
model_id = "RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")
processor.set_num_lookahead_tokens(6) # 0/3/6/13 = 80/320/560/1120 ms
sr = processor.feature_extractor.sampling_rate
audio = load_audio("audio.wav", sampling_rate=sr)
inputs = processor(audio, sampling_rate=sr).to(model.device, dtype=model.dtype)
out = model.generate(**inputs, return_dict_in_generate=True)
print(processor.decode(out.sequences, skip_special_tokens=True))processor.decode returns a list of strings, one per utterance in the batch.
NeMo
apt-get update && apt-get install -y libsndfile1 ffmpeg
pip install "nemo_toolkit[asr]==3.0.0" "numba-cuda[cu12]" "numpy==2.2.6" \
jiwer soundfile huggingface_hubInstallation issues: see the NVIDIA guide.
The NeMo path is prompt-conditioned through the manifest — it reads the language from the lang field, so a bare file path raises ValueError: Unknown prompt key: 'None'. This does not apply to the Transformers path above, where the language is baked into the config.
import json, tempfile, os, torch
from huggingface_hub import hf_hub_download
from nemo.collections.asr.models import ASRModel
CKPT = hf_hub_download(
"RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian",
"nemotron35-jordanian.nemo",
)
model = ASRModel.restore_from(CKPT, map_location="cuda")
model.eval()
model.encoder.set_default_att_context_size([56, 6]) # 560 ms chunks
def transcribe(model, wav_paths, batch_size=8):
rows = [{"audio_filepath": p, "duration": 0.0, "text": "", "lang": "ar"}
for p in wav_paths]
with tempfile.NamedTemporaryFile("w", suffix=".json", delete=False,
encoding="utf-8") as tmp:
for r in rows:
tmp.write(json.dumps(r, ensure_ascii=False) + "\n")
path = tmp.name
try:
with torch.inference_mode():
out = model.transcribe(path, batch_size=batch_size)
finally:
os.unlink(path)
if isinstance(out, tuple):
out = out[0]
return [h.text if hasattr(h, "text") else str(h) for h in out]
print(transcribe(model, ["sample.wav"], batch_size=1)[0])Evaluate on a manifest
One JSON object per line with audio_filepath, duration, text, lang. duration is in seconds; lang must be "ar" on every row.
{"audio_filepath": "/data/test/utt_0001.wav", "duration": 3.42, "text": "أهلا وسهلا كيف بقدر أساعدك", "lang": "ar"}
{"audio_filepath": "/data/test/utt_0002.wav", "duration": 5.18, "text": "بدي أستفسر عن الطلبية اللي عملتها مبارح", "lang": "ar"}
{"audio_filepath": "/data/test/utt_0003.wav", "duration": 1.07, "text": "تفضل", "lang": "ar"}import json, time
from nemo.collections.asr.metrics.wer import word_error_rate
rows = [json.loads(l) for l in open("test_manifest.json", encoding="utf-8") if l.strip()]
refs = [r["text"] for r in rows]
t0 = time.time()
hyps = transcribe(model, [r["audio_filepath"] for r in rows], batch_size=8)
elapsed = time.time() - t0
audio_s = sum(r["duration"] for r in rows)
print(f"WER {word_error_rate(hypotheses=hyps, references=refs)*100:.2f}%")
print(f"CER {word_error_rate(hypotheses=hyps, references=refs, use_cer=True)*100:.2f}%")
print(f"RTF {elapsed/audio_s:.4f}")For Arabic, apply the same normalisation to references and hypotheses before scoring — folding diacritics, alef/ya/ta-marbuta variants and punctuation — or the number is not comparable to anyone else's. Keep the normaliser's Arabic character classes written as \uXXXX escapes: as literal Arabic, bidirectional reordering in an editor silently swaps the endpoints of a character range, and the normaliser then deletes the text it was meant to fold.
Streaming
python examples/asr/asr_cache_aware_streaming/speech_to_text_cache_aware_streaming_infer.py \
model_path=/path/to/nemotron35-jordanian.nemo \
dataset_manifest=test_manifest.json \
batch_size=1 \
target_lang=ar \
att_context_size="[56,6]" \
strip_lang_tags=true \
output_path=./stream_outSecond value of att_context_size sets latency: 0 = 80 ms, 1 = 160 ms, 3 = 320 ms, 6 = 560 ms, 13 = 1120 ms. Left context stays at 56 frames.
Hardcode ar. Do not use target_lang=auto — it reintroduces the tag emission this fine-tune eliminated. Keep strip_lang_tags=true as insurance on out-of-distribution audio.
Scope and intended use
Built for voice agents: intent classification, slot filling, keyword spotting — tasks where a downstream language model absorbs word-level errors and the 27% relative gain translates directly into better intent accuracy.
Numeric slots need a conversion pass. Numbers come out as Arabic words, not digits, so phone numbers, order IDs and amounts require word-to-number post-processing. See Numbers are written as words.
Not for verbatim transcription. At 28.8% WER, output should not be read unedited or used for compliance recording.
Code-switching is the hardest category at 36.45% WER (n=99), though also where fine-tuning helped most — the base scored 56.87%. Single borrowed words transliterate into Arabic (ال manager → المانيجر), which is arguably correct but scores as a substitution. Continuous English degrades further. Worth knowing if your calls open with a scripted English greeting.
Corpus-level gain, not universal. 226 of 1,311 utterances regressed after fine-tuning; the headline number is a large improvement on most, net of real regression on some.
Domain scope. Tuned for Jordanian dialect on this corpus's channel conditions and transcription conventions. The test set comes from a different sector than the training data, with different speakers and topics — so in-domain performance is likely better than these numbers suggest. Performance on MSA or other Arabic dialects is unmeasured.
Training
Method: full fine-tune, all 637M parameters updated, none frozen. NeMo has no first-class LoRA path for FastConformer-RNNT, and the model fits on one L40S at this batch size. Tokenizer kept from the base checkpoint.
Validation carved out of train and split by recording ID — a per-utterance split would leak speaker and session across the boundary. Test never trained or validated on.
Hyperparameters:
Hardware: single NVIDIA L40S. Python 3.12.6, torch 2.8.0+cu129, NeMo 3.0.0, numpy 2.2.6. 66.9 minutes for 10 epochs ≈ 1.1 GPU-hours.
Fine-tuning this base model yourself — three config traps
optim.lris a Noam multiplier, not a learning rate: peak = lr × 1.40e-3. The default 2.0 gives 6.3e-4, too hot for a warm checkpoint; a sane-looking 1e-4 gives 1.4e-7 and trains on nothing.optim.sched.d_modelmust be a literal1024. The YAML interpolation is detached by the time Lightning callsconfigure_optimizers→InterpolationKeyError.use_bucketingdefaults to false, sonum_bucketsis silently ignored.is_tarredmust be false for loose.wavfiles.
Training data
Trained on a private Jordanian dialect Arabic corpus, not publicly available. The test set is private, so the numbers here are reported for base-vs-fine-tuned comparison under identical conditions rather than as an independently reproducible benchmark.
License
Governed by OpenMDW-1.1, inherited from nvidia/nemotron-3.5-asr-streaming-0.6b.
Citation
@misc{nemotron35_jordanian,
title = {Nemotron 3.5 ASR fine-tuned for Jordanian Dialect Arabic},
author = {FILL IN},
year = {2026},
howpublished = {\url{https://huggingface.co/RamaBashar22/nemotron-3.5-asr-streaming-0.6b-jordanian}}
}Base model:
@misc{nemotron35asr,
title = {Nemotron 3.5 ASR},
author = {NVIDIA},
year = {2026},
howpublished = {\url{https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b}}
}References: Stateful Conformer with Cache-based Inference · Fast Conformer
Maintainer: FILL IN · Questions and issues via the Community tab.
