datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
legal2023_38hrs
legal2023_38hrs
Court-audio ASR dataset: 38.6 h of English legal/court speech cut into
per-speaker segments, with speaker-disjoint train / validation / test splits.
⚠️ Pseudo-labels, not gold. Transcripts are produced by an automatic pipeline
not human annotation. Corpus WER vs an independent judge (nvidia/parakeet-rnnt-1.1b) is ~20%.
A per-segment confidence avg_score is provided; only segments with avg_score >= 0.4
are included. Filter further on segment_wer if you need… See the full description on the dataset page: https://huggingface.co/datasets/vanarp/legal2023_38hrs.in22-legal
IN22-Legal
Test-only out-of-distribution legal-domain dictation benchmark for Indic ASR. Read-speech recordings of legal passages from the IN22-Gen corpus, dense in domain entities (statute names, section numbers), formal numerals (dates, monetary amounts), and complex clause structures.
Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (Interspeech 2026, under review).
📄 Documentation: DATASHEET.md… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/in22-legal.Legal_audio_dataset_aws_free_code_camp_trimmed3Legal_audio_dataset_aws_free_code_campLegal_audio_dataset_tedx1Legal_audio_dataset1Legal_audio_dataset_aws_free_code_camp_trimmedLegal_audio_dataset333Legal_audio_dataset_tedxLegal_audio_dataset3Legal_audio_dataset_whisper_ttslegal-services-call-audio-redacted-org1
Redacted Legal Services Call Audio - Org 1
This is a private, proprietary dataset of redacted legal-services call audio
and aligned transcripts. The dataset is intended for authorized use by approved users,
customers, and organizations under applicable commercial agreements.
Contents
Rows: 559
Split: train
Audio format: lossless FLAC converted from redacted WAV audio
Shards: 8 WebDataset-style TAR files under shards/
Viewer table: data/train.parquet… See the full description on the dataset page: https://huggingface.co/datasets/sumo-torii/legal-services-call-audio-redacted-org1.
