datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
legal2023_38hrs
legal2023_38hrs
Court-audio ASR dataset: 38.6 h of English legal/court speech cut into
per-speaker segments, with speaker-disjoint train / validation / test splits.
⚠️ Pseudo-labels, not gold. Transcripts are produced by an automatic pipeline
not human annotation. Corpus WER vs an independent judge (nvidia/parakeet-rnnt-1.1b) is ~20%.
A per-segment confidence avg_score is provided; only segments with avg_score >= 0.4
are included. Filter further on segment_wer if you need… See the full description on the dataset page: https://huggingface.co/datasets/vanarp/legal2023_38hrs.in22-legal
IN22-Legal
Test-only out-of-distribution legal-domain dictation benchmark for Indic ASR. Read-speech recordings of legal passages from the IN22-Gen corpus, dense in domain entities (statute names, section numbers), formal numerals (dates, monetary amounts), and complex clause structures.
Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (Interspeech 2026, under review).
📄 Documentation: DATASHEET.md… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/in22-legal.
