CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sambal /speech_400kaudio100K<n<1M0 likes396 downloads2y agoHugging Face02sambal /speech_dataaudio1M<n<10M0 likes130 downloads2y agoHugging Face03Reord-AI /english-casual-speech-sample-south-african-accentgated English Casual Speech Sample (South African Accent) South African crowd-sourced participants respond to questions about their daily lives and activities. This dataset is a sample of a larger collection from the same data collection campaign. Changelog FEB 2026: initial share. ASR (Chirp3) transcripts. WER: 12% Specs Speakers: ~550 unique South African speakers Total duration: ~60 hours Files sample rate: 48kHz Actual sample rate: TBD Language: English (SA… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/english-casual-speech-sample-south-african-accent.audio10K<n<100K0 likes49 downloads8mo agoHugging Face04DavidErikMollberg /samromur_asr Dataset Card for samromur_asr Dataset Summary This is a modfied copy of the dataset from The Language and Voice Laboratory in RU. This is the first release of the Samrómur Icelandic Speech corpus that contains 100.000 validated utterances. The corpus is a result of the crowd-sourcing effort run by the Language and Voice Lab at the Reykjavik University, in cooperation with Almannarómur, Center for Language Technology. Languages The audio is in Icelandic. The… See the full description on the dataset page: https://huggingface.co/datasets/DavidErikMollberg/samromur_asr.audioautomatic-speech-recognition100K<n<1M0 likes34 downloads1y agoHugging Face05sambal /speech_40kaudio100K<n<1M0 likes21 downloads2y agoHugging Face06sambal /arabic_speech_data_8.taraudio10K<n<100K1 likes15 downloads2y agoHugging Face07sam8000 /EuroSpeech-WebDatasetaudio1K<n<10K0 likes12 downloads1y agoHugging Face08krishnakalyan3 /vb_sampleaudion<1K0 likes8 downloads1y agoHugging Face09scotus-sim /scotus-samuel_a_alito_jr-audio SCOTUS-sim audio: samuel_a_alito_jr Per-utterance audio clips from Oyez oral-argument mp3s, sliced at the start_time / stop_time timestamps stored in the companion scotus-sim/scotus-samuel_a_alito_jr-training dataset. Alignment clip_NNNNN.wav in the tarball corresponds exactly to audio_segments.jsonl[NNNNN] in the training companion dataset. In metadata.jsonl each row carries the same 0-padded index in idx. This supersedes the v1 tarball, which had systematic… See the full description on the dataset page: https://huggingface.co/datasets/scotus-sim/scotus-samuel_a_alito_jr-audio.audio10K<n<100K0 likes8 downloads5mo agoHugging Face10otoearth /otoSpeech-HQ-full-duplex-samplesgated Dataset Card for otoSpeech-HQ-full-duplex-samples: Full-Duplex Conversational Speech Dataset Samples Dataset Summary otoSpeech-HQ-full-duplex-samples is a curated collection of high-quality full-duplex conversational speech samples designed for commercial and production-oriented use. This repository is derived from a private subset of otoSpeech and features carefully selected English two-speaker conversations with enhanced audio quality. The samples are intended for… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-HQ-full-duplex-samples.audioaudio-to-audion<1K1 likes7 downloads6mo agoHugging Face11krishnakalyan3 /emo_speech_sampleaudio1K<n<10K1 likes6 downloads2y agoHugging Face12sambal /lava_datasetaudio100K<n<1M0 likes6 downloads2y agoHugging Face13zello /zello-public-channels-voice-samplegated Zello Public Channels Voice Dataset Sample Dataset summary Total audio: 48.23 hours Total messages: 13207 Breakdown by language Language Hours Messages Speakers Channels ms 10.65 2056 112 7 en 9.54 2990 362 26 id 6.98 1597 157 11 es 6.86 2144 224 21 tl 4.28 1404 78 6 pt 3.92 863 66 11 sw 2.43 182 21 1 ru 1.30 282 42 12 th 0.95 55090 3 zh 0.87 828 312 10 it 0.15 25 19 7 is 0.09 76 25 10 ko 0.05 75 43 13 vi 0.04 42 34 8 fr… See the full description on the dataset page: https://huggingface.co/datasets/zello/zello-public-channels-voice-sample.audioaudio-to-audio10K<n<100K1 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.