CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vanarp /legal2023_38hrs legal2023_38hrs Court-audio ASR dataset: 38.6 h of English legal/court speech cut into per-speaker segments, with speaker-disjoint train / validation / test splits. ⚠️ Pseudo-labels, not gold. Transcripts are produced by an automatic pipeline not human annotation. Corpus WER vs an independent judge (nvidia/parakeet-rnnt-1.1b) is ~20%. A per-segment confidence avg_score is provided; only segments with avg_score >= 0.4 are included. Filter further on segment_wer if you need… See the full description on the dataset page: https://huggingface.co/datasets/vanarp/legal2023_38hrs.audioautomatic-speech-recognition10K<n<100K0 likes138 downloads3mo agoHugging Face02adalat-ai /in22-legal IN22-Legal Test-only out-of-distribution legal-domain dictation benchmark for Indic ASR. Read-speech recordings of legal passages from the IN22-Gen corpus, dense in domain entities (statute names, section numbers), formal numerals (dates, monetary amounts), and complex clause structures. Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (Interspeech 2026, under review). 📄 Documentation: DATASHEET.md… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/in22-legal.audioautomatic-speech-recognitionn<1K0 likes49 downloads4mo agoHugging Face03Sarvesh2003 /Legal_audio_dataset_aws_free_code_camp_trimmed3audio10K<n<100K0 likes22 downloads2y agoHugging Face04Sarvesh2003 /Legal_audio_dataset_aws_free_code_campaudio1K<n<10K0 likes21 downloads2y agoHugging Face05Sarvesh2003 /Legal_audio_dataset_tedx1audion<1K0 likes20 downloads2y agoHugging Face06Sarvesh2003 /Legal_audio_dataset1audion<1K0 likes19 downloads2y agoHugging Face07Sarvesh2003 /Legal_audio_dataset_aws_free_code_camp_trimmedaudion<1K0 likes15 downloads2y agoHugging Face08Sarvesh2003 /Legal_audio_dataset333audio10K<n<100K0 likes11 downloads2y agoHugging Face09Sarvesh2003 /Legal_audio_dataset_tedxaudion<1K0 likes9 downloads2y agoHugging Face10Sarvesh2003 /Legal_audio_dataset3audion<1K0 likes4 downloads2y agoHugging Face11Sarvesh2003 /Legal_audio_dataset_whisper_ttsaudio10K<n<100K0 likes4 downloads2y agoHugging Face12sumo-torii /legal-services-call-audio-redacted-org1gated Redacted Legal Services Call Audio - Org 1 This is a private, proprietary dataset of redacted legal-services call audio and aligned transcripts. The dataset is intended for authorized use by approved users, customers, and organizations under applicable commercial agreements. Contents Rows: 559 Split: train Audio format: lossless FLAC converted from redacted WAV audio Shards: 8 WebDataset-style TAR files under shards/ Viewer table: data/train.parquet… See the full description on the dataset page: https://huggingface.co/datasets/sumo-torii/legal-services-call-audio-redacted-org1.tabularn<1K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.