CoolFace
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pengyizhou /IALP-2026-data IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC Adaptation." This repository holds the fixed target-query, validation, and evaluation sets used across all experiments. Each part is a self-contained .tar.gz. All audio is 16 kHz mono. Each split ships with: audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16) wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.audioautomatic-speech-recognition10K<n<100K0 likes21 downloads3mo agoHugging Face02pengyizhou /hub5_english_eval_2000_swb1gatedaudioautomatic-speech-recognition1K<n<10K0 likes3 downloads1y agoHugging Face03pengyizhou /seamless-interaction-audio-only Seamless Interaction — audio only Audio-only repackaging of Meta's Seamless Interaction dataset (improvised and naturalistic subsets), CC-BY-NC 4.0. Changes from the original: video and motion (npz) data removed; audio resampled from 48 kHz float to 24 kHz 16-bit PCM, stored as FLAC. All label fields are kept as published: vad, transcript: the per-file metadata/*.jsonl records. annotation_1P_IS, annotation_1P_R, annotation_3P_IS, annotation_3P_R, annotation_3P_V: the raw… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/seamless-interaction-audio-only.audioautomatic-speech-recognition100K<n<1M0 likes11m agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.