CoolFace
Datasetpublic

vanarp/legal2023_38hrs

legal2023_38hrs Court-audio ASR dataset: 38.6 h of English legal/court speech cut into per-speaker segments, with speaker-disjoint train / validation / test splits. ⚠️ Pseudo-labels, not gold. Transcripts are produced by an automatic pipeline not human annotation. Corpus WER vs an independent judge (nvidia/parakeet-rnnt-1.1b) is ~20%. A per-segment confidence avg_score is provided; only segments with avg_score >= 0.4 are included. Filter further on segment_wer if you need… See the full description on the dataset page: https://huggingface.co/datasets/vanarp/legal2023_38hrs.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes127downloads
settings

This repository belongs to vanarp on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namelegal2023_38hrs
visibilitypublic
licenceother
gatedno
ownervanarp
Account settings
vanarp/legal2023_38hrs · CoolFace