CoolFace
Datasetpublic

kirandevraj/supreme-court-hearings-asr

Indian Supreme Court Hearings — ASR dataset Sentence-level, force-aligned audio–text pairs from Indian Supreme Court hearings, prepared for fine-tuning ASR models (e.g. Whisper). 46.9 hours across 23 hearings / 15 cases. Load from datasets import load_dataset ds = load_dataset("kirandevraj/supreme-court-hearings-asr") ds["test"][0] # {'audio': {'array', 'sampling_rate': 16000}, 'text': '...', ...} Splits split clips hours train 24,422… See the full description on the dataset page: https://huggingface.co/datasets/kirandevraj/supreme-court-hearings-asr.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes110downloads
Dataset Card

Indian Supreme Court Hearings — ASR dataset

Sentence-level, force-aligned audio–text pairs from Indian Supreme Court hearings, prepared for fine-tuning ASR models (e.g. Whisper). 46.9 hours across 23 hearings / 15 cases.

Load

python
from datasets import load_dataset
ds = load_dataset("kirandevraj/supreme-court-hearings-asr")
ds["test"][0]   # {'audio': {'array', 'sampling_rate': 16000}, 'text': '...', ...}

Splits

splitclipshours
train24,42237.44
validation1,9663.32
test3,7996.11
total30,18746.87

Split is by case (all hearing-dates of a case go to one split) → no speaker/topic leakage.

Fields

  • —audio — 16 kHz mono WAV, decoded on access
  • —text — reference transcript (cased + punctuated)
  • —hearing_id, speaker, start, end, duration
  • —winner — which aligner's boundary was chosen (mms / whisper / mms_only / whisper_only)
  • —chunked — True if the clip was carved from a split (>28 s) sentence

Gold config — human-verified eval set

A separate gold config holds 109 manually reviewed & corrected clips (a hand-checked subset; ~89% of the reviewed sample were already correctly aligned, the rest fixed or dropped) — the trustworthy anchor for evaluation.

python
gold = load_dataset("kirandevraj/supreme-court-hearings-asr", "gold", split="test")

How it was built

  • —Transcripts extracted from PDFs (pdftotext, no OCR), cleaned + sentence-segmented (pysbd).
  • —Two forced alignments per sentence: torchaudio MMS_FA (CTC + <star>) and Whisper large-v3 (stable-ts cross-attention); long files via faster-whisper anchoring / windowing.
  • —Best-of-both selection: each sentence is cut both ways, transcribed by an independent judge (whisper-medium), and the lower-CER boundary is kept (MMS won ~88%, Whisper ~12% — see the winner column).
  • —Sentences where both alignments fail (min CER > 50%) are dropped as misaligned; RMS silence cleaning; chunked to ≤28 s.
  • —Post-chunk CER filter: because that misalignment filter is per sentence (before chunking), each chunk is re-judged by whisper-medium and those over CER 50 are dropped from the test split (51 mis-cut clips removed); every clip carries a chunked flag. A separate human-verified gold config anchors quality.

License / source

Derived from publicly available Supreme Court of India hearing recordings and their official transcripts. Verify applicable licensing/terms before redistribution or commercial use.