CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K6 likes6.6k downloads7mo agoHugging Face02pollen-robotics /microduck-emotions Microduck Emotions A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.audioroboticsn<1K6 likes898 downloads17d agoHugging Face03eturok-weizmann /laser-vibrations Laser Vibrations Dataset of laser speckle vibration recordings used to locate objects hidden inside a cardboard box. A 10×10 grid of lasers shines on the side of a box; as loudspeakers excite the box, each laser's speckle pattern shifts in proportion to the local surface vibration. The goal is to reconstruct the shape and location of an object inside the box from the vibration signals alone. Per-sample viewer metadata lives in data/metadata.jsonl. Full signal data and media files… See the full description on the dataset page: https://huggingface.co/datasets/eturok-weizmann/laser-vibrations.audion<1K0 likes725 downloads4mo agoHugging Face04ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes613 downloads3mo agoHugging Face05yuanzhuyun /asr-reference-set-eval-temp Temporary ASR evaluation audio Temporary public audio files used for hosted ASR evaluation. audio1K<n<10K0 likes356 downloads2mo agoHugging Face06MrJackTung /cs-envi-dual-encoder-60audion<1K0 likes237 downloads4mo agoHugging Face07m-a-p /EMOaudion<1K1 likes214 downloads1y agoHugging Face08vhands /audio-event-classification-post-public audio-event-classification-post-public Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.textaudio-classification100K<n<1M1 likes181 downloads3mo agoHugging Face09MohamedGomaa30 /EGYSpeak EGYSpeak A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline. Quick Start 1. Download the dataset: from huggingface_hub import snapshot_download snapshot_download( repo_id="MohamedGomaa30/EGYSpeak", repo_type="dataset", local_dir="EGYSpeak", ) 2. Extract the dataset: from… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/EGYSpeak.textautomatic-speech-recognition100K<n<1M1 likes158 downloads5mo agoHugging Face10EnjunDu /OST-Diagbenchgated OST-DiagBench OST-DiagBench contains exactly four diagnostic operations with 510 cases per operation: 2,040 cases in total. Every case includes an MP4 input and a paired M4A audio file. Academic research use only, no commercial use. The clips are short, transformed excerpts derived from public academic audio-visual corpora whose material originates from user-uploaded YouTube content; the donor sounds come from ESC-50 under CC BY-NC 3.0. We are not the rights holders of the… See the full description on the dataset page: https://huggingface.co/datasets/EnjunDu/OST-Diagbench.audiovisual-question-answering1K<n<10K6 likes129 downloads14d agoHugging Face11hojreh /hawza-asr-evalA test dataset for evaluate ASR (Automatic Speech Recognition) models in the domain of Islamic lectures and specialized Hawza courses. Audio files are mono 16khz wav. Texts are verified. audion<1K1 likes129 downloads20d agoHugging Face12playwithmino /alimeeting-eval-8k AliMeeting Eval — 8 kHz CH0 clips Unknown-(N) eval clips from AliMeeting Eval (M2MeT / OpenSLR 119), far-field channel 0, resampled to 8 kHz. Manifest Clips manifests/eval_all.jsonl 1280 (full local eval) manifests/eval_n100.jsonl 100 stratified subset manifests/eval_n200.jsonl 200 stratified subset manifests/eval_*mix.jsonl by (N=1\ldots4) Headset s{k}.wav stems (near, TextGrid-gated) are present for the n200 subset (eval_n200_headset.jsonl). Other clips… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/alimeeting-eval-8k.audioaudio-to-audio1K<n<10K0 likes126 downloads23d agoHugging Face13uzinfocom-edu-ai /uzbek-asr-curated-701h Uzbek ASR Curated Dataset (701 hours) A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation. Dataset Description Language Uzbek (Latin script with okina ʻ) Total utterances 337,920 Total duration ~701 hours Audio format 16 kHz mono WAV (PCM_16) Manifest format NeMo JSONL Splits train (94%) / val (3%) / test (3%) Splits Split Utterances Hours Train 317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.audioautomatic-speech-recognition100K<n<1M1 likes111 downloads3mo agoHugging Face14eQOURSE /multilingual-speech Multilingual Indian Conversational Speech A dataset of naturalistic, spontaneous two-speaker conversations across 13 Indian languages, with segment-level transcripts, speaker profiles, timestamps, and recording metadata. Designed for ASR, TTS, speaker diarization, and conversational speech research. Languages (13) Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali, Odia, Punjabi, Tamil, Telugu, Urdu. Content Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.audioautomatic-speech-recognitionn<1K0 likes88 downloads3mo agoHugging Face15arcada-labs /event-bench Event Bench 29-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as an event planning assistant. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as an event planning assistant managing venue bookings, catering, and guest logistics. The conversation features cascading changes — a venue switch triggers catering… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/event-bench.audioautomatic-speech-recognitionn<1K2 likes77 downloads6mo agoHugging Face16vocence /vocence_eval_corpus Dataset Card for vocence_eval_corpus Dataset Summary Everything produced by evaluating 7 PromptTTS systems (gemini, voxcpm, qwen3, maya1, openai, parler-tts, elevenlabs) on the 200-item balanced subset of vocence_corpus: the generated audio, every per-clip judge output, and the aggregated leaderboards/charts. For the prompt corpus alone (no eval data), see vocence_corpus. Scoring library: vocencebench (source) Benchmark code + full methodology write-up:… See the full description on the dataset page: https://huggingface.co/datasets/vocence/vocence_eval_corpus.audiotext-to-speechn<1K0 likes71 downloads2mo agoHugging Face17Noothi /telugu-indicf5-evaluationaudion<1K0 likes65 downloads2mo agoHugging Face18danielrosehill /English-Hebrew-Mixed-Sentences English-Hebrew Mixed Sentences Dataset A dataset of English sentences with Hebrew words and phrases interspersed, designed for speech-to-text training and evaluation for English speakers in Israel. Overview This dataset addresses a common challenge for English-speaking immigrants in Israel: standard speech-to-text (STT) systems struggle to accurately transcribe code-switched speech where Hebrew words are mixed into primarily English sentences. Example: "I need to pick up… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/English-Hebrew-Mixed-Sentences.audion<1K0 likes62 downloads10mo agoHugging Face19jordand /echo-embeddings-vctk-tar VCTK Speaker Embeddings (tarred) Items: 109 This dataset ships as a single tar at the repo root. Members preserve paths like VCTK/<id>/audio.mp3 and VCTK/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the CSTR VCTK Corpus. Distributed under CC BY 4.0; attribution required. textn<1K0 likes59 downloads10mo agoHugging Face20jordand /echo-embeddings-expresso-tar Expresso Speaker Embeddings (tarred) Items: 17 This dataset ships as a single tar at the repo root. Members preserve paths like Expresso/<id>/audio.mp3 and Expresso/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the Expresso dataset (INTERSPEECH 2023). Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted. textn<1K0 likes58 downloads10mo agoHugging Face21jordand /echo-embeddings-custom Custom Speaker Embeddings Contains speaker folders within HF-Custom, each with: a precomputed speaker embedding (speaker_latent.safetensors) its corresponding audio (audio.mp3) a metadata file describing the voice and licensing (metadata.json) Licensing:There is no single license for this dataset. Each voice has its own terms stored in its metadata.json. You must check the metadata for any voice you use. audion<1K3 likes58 downloads10mo agoHugging Face22ODYSSEYAILABS /odyssey-v0.5.1-16h-eval-preview-public 🇿🇦 Odyssey V0.5.1 — 16H Evaluation Preview (Code-Switched ASR) Odyssey AI Labs presents a 16-hour Enterprise Evaluation Preview of the V0.5 corpus. This release is designed to surface the hardest South African ASR realities directly: native urban code-switching, multi-speaker turn-taking, and overlapping conversational speech. A structured public preview of the Odyssey SA Voice Corpus, designed for researchers, data buyers, and speech teams evaluating multilingual South African… See the full description on the dataset page: https://huggingface.co/datasets/ODYSSEYAILABS/odyssey-v0.5.1-16h-eval-preview-public.tabularn<1K0 likes58 downloads5mo agoHugging Face23jordand /echo-embeddings-ears-tar EARS Speaker Embeddings (tarred) Items: 2568 This dataset ships as a single tar at the repo root. Members preserve paths like EARS/<id>/audio.mp3 and EARS/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the EARS dataset. Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted. text1K<n<10K1 likes52 downloads10mo agoHugging Face24Rakancorle1 /thud-eval THUD-Eval · audio-visual Clever Hans benchmark Evaluation benchmark accompanying the paper When Vision Speaks for Sound. This dataset probes the audio-visual Clever Hans effect — the tendency of video-capable MLLMs to appear to listen while really just reading visual cues. We test the same source clips under three audio interventions: Task Intervention What it tests sync audio temporally shifted (early / delay) Can the model detect a time offset? mute audio replaced… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/thud-eval.audion<1K0 likes41 downloads4mo agoHugging Face25mahir2111 /EGYSpeak EGYSpeak A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline. Quick Start 1. Download the dataset: from huggingface_hub import snapshot_download snapshot_download( repo_id="MohamedGomaa30/EGYSpeak", repo_type="dataset", local_dir="EGYSpeak", ) 2. Extract the dataset: from… See the full description on the dataset page: https://huggingface.co/datasets/mahir2111/EGYSpeak.textautomatic-speech-recognition100K<n<1M0 likes39 downloads2mo agoHugging Face26speech-uk /asr-evaluationstabularautomatic-speech-recognition10K<n<100K0 likes38 downloads2y agoHugging Face27Aakash22134 /hi-en-noisy-vad-benchmark Hindi-English Noisy VAD Benchmark Version 0.1.0 is a deterministic, evaluation-only benchmark with 78 mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB). Intended use Use this dataset to compare voice-activity detectors under matched Hindi/English noise conditions and to tune thresholds. It is too small and insufficiently diverse for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.audioaudio-classificationn<1K0 likes37 downloads1mo agoHugging Face28bnovikov /gemma-4-e4b-audio-qa Gemma-4 E4B Audio-QA Training Mix A 91k-row audio question-answering dataset assembled from four public upstream datasets, formatted as ChatML-style conversations for instruction-tuning an audio-language model. This is the exact training data used for bnovikov/gemma-4-e4b-audio-v3. Important: this repository contains only the metadata and prompts/answers. The audio files are NOT hosted here. Each audio_path is a source-tagged ID like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.textaudio-classification10K<n<100K0 likes36 downloads5mo agoHugging Face29primal-sage /emotion-preserving-s2st-benchmark Emotion-Preserving Speech-to-Speech Translation Benchmark Overview First benchmark for evaluating emotion preservation in speech-to-speech translation systems. Pipeline: English speech → Emotion Detection → Translation (EN→HI) → Emotion-conditioned TTS. Key Results (1440 samples, RAVDESS dataset) Metric Score Emotion Detection Accuracy 36.3% (4-class on 8-class data) Emotion Preservation Rate 43.5% Preservation Gain over Flat TTS +7.2% F0… See the full description on the dataset page: https://huggingface.co/datasets/primal-sage/emotion-preserving-s2st-benchmark.audioaudio-classificationn<1K0 likes35 downloads7mo agoHugging Face30MrlolDev /voxtral-emotion-temporal VoxTral Emotion Temporal Dataset Dataset for training emotion transition detection in speech. ~500 clips with frame-level emotion annotations at 20ms resolution. Overview Property Value Clips ~500 Sample Rate 16000 Hz Frame Resolution 20ms Emotions 6 (neutral, happy, angry, sad, surprise, fear) Clip Distribution Single Emotion (40%): One emotion throughout Transitions (40%): Emotion changes mid-sentence (2-3 segments) No Emotion (20%):… See the full description on the dataset page: https://huggingface.co/datasets/MrlolDev/voxtral-emotion-temporal.audion<1K0 likes35 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.