CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K6 likes6.8k downloads7mo agoHugging Face02thanhnew2001 /VietSuperSpeech VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.audio10K<n<100K6 likes3.4k downloads7mo agoHugging Face03zhifeixie /StreamAudio-2M StreamAudio-2M Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips are organised into six task subsets. Subsets Subset Rows Description Stream_Audio_Understanding 90,738 Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA Real_time_ASR 28,109 Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.tabularaudio-classification100K<n<1M30 likes2.8k downloads4mo agoHugging Face04QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads25d agoHugging Face05cm2435-new /gdpval_preference_rubricsaudion<1K0 likes1.6k downloads5mo agoHugging Face06YomnaGharib /dahih-tts2-demucs-cleanedaudio10K<n<100K1 likes1.6k downloads4mo agoHugging Face07projecti7 /phoneme2audio100K<n<1M0 likes977 downloads6mo agoHugging Face08ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes608 downloads3mo agoHugging Face09Harland /DCASE2026-Task5-DevSet DCASE 2026 Task 5 Audio-Dependent Question Answering (ADQA) Development Set This is the official Development Set for DCASE 2026 Challenge Task 5: Audio-Dependent Question Answering (ADQA). The ADQA task focuses on addressing "Textual Hallucination" in Large Audio-Language Models (LALMs) — where models pass audio understanding benchmarks by relying on text prompts and internal linguistic priors rather than actual audio perception. ADQA introduces a rigorous evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Harland/DCASE2026-Task5-DevSet.audio1K<n<10K1 likes566 downloads2mo agoHugging Face10anonymous2222 /Sympatheia-18k Sympatheia-18k Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research. It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by synthesized audio and text transcripts. Dataset Structure Subset Unique Queries Responses Description Emotional 8,400 train / 3,600 eval 8,400 train / 3,600 eval Emotional queries with emotionally-matched responses Neutral 350 train / 150 eval 4,200 train /… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2222/Sympatheia-18k.audioaudio-to-audio10K<n<100K0 likes477 downloads5mo agoHugging Face11taohu /music2chords_v2audio1K<n<10K0 likes362 downloads7mo agoHugging Face12KishoOoOo /comfyui-wan22-assetsaudion<1K1 likes287 downloads25d agoHugging Face13Sheeba2026 /bharatvani-hindi-speech-corpusgated BharatVani Hindi Speech Corpus (150-Hour Studio Dataset) Proprietary Speech Asset • TheCreatorOS • BharatVani AI 1. Overview The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi. Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.audiotext-to-speech100K<n<1M1 likes284 downloads5d agoHugging Face14nips26 /VoxSafeBench VoxSafeBench This dataset is uploaded as raw files (JSONL + audio), not parquet. Subset: Safety-tier1 Split: No_jailbreak (8708 samples) Columns: system_prompt, clean_audio_file_name, diverse_audio_file_name, transcript, super_category, task_type, language, query Split: Singleturn_jailbreak (2516 samples) Columns: system_prompt, audio_file_name, transcript, source_text, super_category, jailbreak_type, task_type, language, query… See the full description on the dataset page: https://huggingface.co/datasets/nips26/VoxSafeBench.audio1K<n<10K1 likes238 downloads5mo agoHugging Face15inesriahi /valor32k-avqa-v2 Valor32k-AVQA v2.0 Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position. Links Paper: ACM Digital Library Project page: inesriahi.github.io/valor32k-avqa-2 Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.tabularquestion-answering100K<n<1M0 likes214 downloads3mo agoHugging Face16joshruby /nightjar-flight-20260814 nightjar-flight-20260814 Elevation datum (fit_el_datum, 2026-08-16) Status: FEW_ON_DRONE — this day's elevation datum is honestly UNSOLVABLE from banked data (camera never/rarely locked on the drone). Close-range elevation truth for this day remains datum-limited (~25 deg floor). Tool: sirch613/subhunt-v2 v3/fit_el_datum.py; summary: joshruby/acoustic-knowledge -> v3/assets/el_datum/SUMMARY.json. audion<1K0 likes204 downloads1mo agoHugging Face17FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes175 downloads6mo agoHugging Face18nyuuzyou /OpenGameArt-GPL-2.0 Dataset Card for OpenGameArt-GPL-2.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 2.0 (GPL-2.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All asset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-2.0.audioimage-classificationn<1K0 likes160 downloads1y agoHugging Face19VibroNav /August25_ChickenZucchini Filename structure e.g. chicken_zucchini_speed-10_22g-12cm-shiba_stethoscope_2025-08-04_19.28.10 meaning: material-speed-needle_size-needle_length-needle_type-microphone_type-timestamp Structure Method of puncturing (Dobot) Needle types audion<1K0 likes116 downloads10mo agoHugging Face20playwithmino /aishell1mix-ver2-n100-per-mix AISHELL-1 Mix ver2 — 100 clips per mix This is AISHELL-1 Mix ver2, not ver1. 8 kHz mono test subset: 100 mixtures per speaker count (N=1\ldots5) (50 mix_clean + 50 mix_both each) → 500 clips. Derived from the local aishell1mix_ver2 test SCPs (data/scp/scp_aishell1mix_ver2). Includes mixture + oracle speaker stems and transcripts. Split Count 1mix / 2mix / 3mix / 4mix / 5mix 100 each clean / both 250 each Files manifests/test.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/aishell1mix-ver2-n100-per-mix.audioaudio-to-audion<1K0 likes67 downloads23d agoHugging Face21Sheeba2026 /bharatvani-hindi-showcase BharatVani Hindi Speech Corpus • Public Interactive Showcase 150-Hour Enterprise Devanagari Hindi Speech Corpus & Precomputed Latents Curated & Mastered by BharatVani AI • TheCreatorOS 1. Interactive Dataset Preview This repository is the official public evaluation showcase for the 150-Hour BharatVani Hindi Speech Corpus (103,784 Studio Clips). Use the Dataset Viewer above to play real audio clips, inspect the word-level timestamp alignments, and… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-showcase.audiotext-to-speechn<1K0 likes61 downloads5d agoHugging Face22meandyou200175 /thauaudio10K<n<100K0 likes60 downloads8mo agoHugging Face23anonymous2026082026 /MMAG-Benchmark MMAG: A Multi‑Control Mixed Audio Generation Benchmark MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment. Dataset Structure The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026082026/MMAG-Benchmark.audiotext-to-audio1K<n<10K0 likes47 downloads2mo agoHugging Face24Aakash22134 /hi-en-noisy-vad-benchmark Hindi-English Noisy VAD Benchmark Version 0.1.0 is a deterministic, evaluation-only benchmark with 78 mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB). Intended use Use this dataset to compare voice-activity detectors under matched Hindi/English noise conditions and to tune thresholds. It is too small and insufficiently diverse for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.audioaudio-classificationn<1K0 likes40 downloads1mo agoHugging Face25zhaochenyang20 /Video_AMME_ci Video-AMME CI Video-AMME is a 50-case CI dataset derived from zhaochenyang20/Video_MME_ci. Each example keeps the Video-MME video and moves the question, answer choices, and answer-format instruction into a spoken WAV file. Files data/test.jsonl: metadata and source Video-MME references. audios/*.wav: spoken question/options/instruction. videos/*.mp4: present only when built with --copy-videos. Generation TTS model: fishaudio/s2-pro Max samples requested: 50… See the full description on the dataset page: https://huggingface.co/datasets/zhaochenyang20/Video_AMME_ci.audion<1K0 likes38 downloads5mo agoHugging Face26mahir2111 /EGYSpeak EGYSpeak A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline. Quick Start 1. Download the dataset: from huggingface_hub import snapshot_download snapshot_download( repo_id="MohamedGomaa30/EGYSpeak", repo_type="dataset", local_dir="EGYSpeak", ) 2. Extract the dataset: from… See the full description on the dataset page: https://huggingface.co/datasets/mahir2111/EGYSpeak.textautomatic-speech-recognition100K<n<1M0 likes36 downloads2mo agoHugging Face27triad-26 /TRIAD TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models TRIAD is a diagnostic benchmark for evaluating whether omni-modal models can resolve ambiguity by jointly integrating text, image, and audio. Overview TRIAD targets a failure mode common in multimodal evaluation: a model may appear to solve a multimodal task while actually relying on only one dominant modality. In TRIAD, each example is constructed so that the full tri-modal input is needed… See the full description on the dataset page: https://huggingface.co/datasets/triad-26/TRIAD.audiovisual-question-answeringn<1K1 likes23 downloads5mo agoHugging Face28Letian2003 /stage1a_smoke_data stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh) Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP → frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend: WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT encoder — no raw-audio decoding at train time. 113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.tabularautomatic-speech-recognitionn<1K0 likes22 downloads3mo agoHugging Face29instinct-org /yt2_chunked_tokenizedgated yt2_chunked_48k_tokenized This is a gated Russian tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes16 downloads4mo agoHugging Face30Leon299 /cmi-annotateaudio1K<n<10K0 likes14 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.