kirandevraj/supreme-court-hearings-asr
Indian Supreme Court Hearings — ASR dataset Sentence-level, force-aligned audio–text pairs from Indian Supreme Court hearings, prepared for fine-tuning ASR models (e.g. Whisper). 46.9 hours across 23 hearings / 15 cases. Load from datasets import load_dataset ds = load_dataset("kirandevraj/supreme-court-hearings-asr") ds["test"][0] # {'audio': {'array', 'sampling_rate': 16000}, 'text': '...', ...} Splits split clips hours train 24,422… See the full description on the dataset page: https://huggingface.co/datasets/kirandevraj/supreme-court-hearings-asr.
Indian Supreme Court Hearings — ASR dataset
Sentence-level, force-aligned audio–text pairs from Indian Supreme Court hearings, prepared for fine-tuning ASR models (e.g. Whisper). 46.9 hours across 23 hearings / 15 cases.
Load
from datasets import load_dataset
ds = load_dataset("kirandevraj/supreme-court-hearings-asr")
ds["test"][0] # {'audio': {'array', 'sampling_rate': 16000}, 'text': '...', ...}Splits
Split is by case (all hearing-dates of a case go to one split) → no speaker/topic leakage.
Fields
audio— 16 kHz mono WAV, decoded on accesstext— reference transcript (cased + punctuated)hearing_id,speaker,start,end,durationwinner— which aligner's boundary was chosen (mms/whisper/mms_only/whisper_only)chunked—Trueif the clip was carved from a split (>28 s) sentence
Gold config — human-verified eval set
A separate gold config holds 109 manually reviewed & corrected clips (a hand-checked subset; ~89% of the reviewed sample were already correctly aligned, the rest fixed or dropped) — the trustworthy anchor for evaluation.
gold = load_dataset("kirandevraj/supreme-court-hearings-asr", "gold", split="test")How it was built
- Transcripts extracted from PDFs (
pdftotext, no OCR), cleaned + sentence-segmented (pysbd). - Two forced alignments per sentence: torchaudio MMS_FA (CTC +
<star>) and Whisper large-v3 (stable-ts cross-attention); long files via faster-whisper anchoring / windowing. - Best-of-both selection: each sentence is cut both ways, transcribed by an independent judge (whisper-medium), and the lower-CER boundary is kept (MMS won ~88%, Whisper ~12% — see the
winnercolumn). - Sentences where both alignments fail (min CER > 50%) are dropped as misaligned; RMS silence cleaning; chunked to ≤28 s.
- Post-chunk CER filter: because that misalignment filter is per sentence (before chunking), each chunk is re-judged by whisper-medium and those over CER 50 are dropped from the test split (51 mis-cut clips removed); every clip carries a
chunkedflag. A separate human-verifiedgoldconfig anchors quality.
License / source
Derived from publicly available Supreme Court of India hearing recordings and their official transcripts. Verify applicable licensing/terms before redistribution or commercial use.
