CoolFace
Modelpublic

adalat-ai/whisper-small-ml-curated-reverse-mft-1-1-1

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes51downloads
Model Card

whisper-small-ml-curated-reverse-mft-1-1-1

Malayalam rich-transcription ASR — emits text with punctuation and numerals natively. Released alongside SCRIBE (Interspeech 2026, under review).

Model details

  • —Architecture: Whisper-small (244M, encoder-decoder)
  • —Base model: `openai/whisper-small`
  • —Language: Malayalam (ml)
  • —Output style: rich transcription (punctuation + formatted numerals)
  • —Training: three-stage curriculum fine-tune on LLM-curated rich-transcription data (diversity → pace/style → precision); this is the stage-3 checkpoint.

Evaluation

Evaluated with SCRIBE. Metrics: WER (lexical, sandhi-aware), LER (legal entities), NER (numerals), PER (punctuation), TER = sum of categorical rates.

MetricFLEURS-ROIN22-Legal
WER (lexical)14.77 %19.47 %
LER (legal entities)–1.59 %
NER (numeral)0.59 %1.20 %
PER (punctuation)14.03 %12.96 %
TER29.39 %35.22 %
Sandhi resolutions44975

FLEURS-RO LER (0.02%) is folded into WER — too sparse on a general-domain set to warrant a row. SCRIBE WER ≠ jiwer WER: SCRIBE's WER is the lexical-category rate after sandhi-tolerant alignment, not monolithic edit distance. Paper uses ERlex / ERnum / ERpunc / ERent; this library uses WER / NER / PER / LER for the same quantities. Paper Table 1 reports the whisper-medium counterpart of this checkpoint.

Usage

python
from transformers import pipeline
asr = pipeline(
    "automatic-speech-recognition",
    model="adalat-ai/whisper-small-ml-curated-reverse-mft-1-1-1",
    generate_kwargs={"language": "ml", "task": "transcribe"},
)
print(asr("sample.wav")["text"])

CTranslate2 build: `adalat-ai/ct2-whisper-small-ml-curated-reverse-mft-1-1-1-ct2-fp16`.

Intended use

Malayalam dictation in legal, medical, and classroom settings where rich-transcription output is required.

License

Apache-2.0. Base model openai/whisper-small is MIT.