CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NathanRoll /global-news-radio-debug Global News Radio Dataset (1 hour per station) Every news radio station from the Radio Browser API, recorded for 1 hour each. Attempted 3037 Successful 2553 Failed 484 Total audio 21 hours Parquet shards 256 Size 0.6 GB Format MP3 16kHz mono 64kbps Usage from datasets import load_dataset ds = load_dataset("NathanRoll/global-news-radio-debug", streaming=True) for sample in ds["train"]: print(sample["station_name"], sample["language"]… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-debug.audioautomatic-speech-recognition1K<n<10K0 likes127 downloads6mo agoHugging Face02muhtasham /whisper-non-verbal-debugaudio10K<n<100K1 likes121 downloads1y agoHugging Face03gavinlaw /chinese-lips-longform-debug Chinese-LiPS Long-Form (zh long streaming speech) Reconstructed continuous long-speech streams from BAAI/Chinese-LiPS, for slide-aware / streaming speech-translation development and evaluation. Each source video (one speaker, one scripted lecture with slides) was released as pre-segmented clips; here they are re-joined into the full talk. Two variants of the same 3 talks (~97 min speech total): config how segments are placed use orig_timeline at their original session… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-longform-debug.audioautomatic-speech-recognition1K<n<10K0 likes66 downloads2mo agoHugging Face04totoluo /gigaspeech_debugaudio100K<n<1M0 likes60 downloads2y agoHugging Face05deep9539 /flickr-audio-image-debugaudio10K<n<100K0 likes32 downloads2mo agoHugging Face06argmaxinc /librispeech-debugaudion<1K0 likes26 downloads3y agoHugging Face07argmaxinc /earnings22-debugaudion<1K0 likes18 downloads3y agoHugging Face08argmaxinc /common_voice_17_0-debugaudion<1K0 likes18 downloads2y agoHugging Face09argmaxinc /common_voice_17_0-debug-zipaudion<1K0 likes18 downloads2y agoHugging Face10EduardoPacheco /earnings22-chunked-debugaudion<1K0 likes16 downloads2y agoHugging Face11Alarak /librispeech_debugaudion<1K0 likes12 downloads2y agoHugging Face12o0dimplz0o /Debug-Test-Setaudio1K<n<10K0 likes11 downloads2y agoHugging Face13arda-argmax /fastmss-debug-v0.5 FastMSS synthetic multi-speaker meetings - parquet edition Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring. Subsets and splits debug_v0.5 — splits: train — 10 mixtures, 20.5 min total, 39 unique… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/fastmss-debug-v0.5.audioautomatic-speech-recognitionn<1K0 likes10 downloads4mo agoHugging Face14DebuggerLab /baby_cry_streamingaudion<1K1 likes9 downloads1y agoHugging Face15bakhtiari77 /debug_v1.1audio10K<n<100K0 likes6 downloads1y agoHugging Face16totoluo /wavcaps_debugaudion<1K0 likes5 downloads2y agoHugging Face17bakhtiari77 /debug_v1.2audion<1K0 likes4 downloads1y agoHugging Face18kehanlu /debuggatedaudion<1K0 likes2 downloads2y agoHugging Face19debugzxcv /nana7migatedaudion<1K0 likes1 downloads4y agoHugging Face20o0dimplz0o /Debug-Test-Set-4audio1K<n<10K0 likes2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.