CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01smgjch /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.audio10K<n<100K3 likes1.5k downloads5mo agoHugging Face02parler-tts /mls_eng_10k Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.audioautomatic-speech-recognition1M<n<10M31 likes1.1k downloads2y agoHugging Face03psdn-ai /bangla-10kgated Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh Bangla-10K is a 10,070.8-hour Bengali speech corpus with 567,323 recordings from India and Bangladesh. It combines scripted single-speaker read speech with natural multi-speaker conversations for Bengali automatic speech recognition (ASR). The paper rounds the corpus scale to 10,000 hours. The corpus and its ASR evaluation are described in the anonymous manuscript… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.audioautomatic-speech-recognition100K<n<1M0 likes324 downloads2d agoHugging Face04seonglae /vls-10k VLS 10K 9,987 MS-COCO images paired with a long written description, a one-sentence spoken summary of that description, the spoken audio, and that audio pre-encoded by two neural codecs. Images and audio are embedded in the parquet, so the viewer renders them and one call opens the set: from datasets import load_dataset ds = load_dataset("seonglae/vls-10k", split="train") ds[0]["image"] # PIL image ds[0]["audio"] # decoded waveform ds[0]["sst"] # the sentence that was… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/vls-10k.audiotext-to-speech1K<n<10K0 likes173 downloads29d agoHugging Face05SynDataLab-JA-Refs /Irodori-Ja-Spk4-10k SynDataLab/Irodori-Ja-Spk4-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk4): 30s female, news-anchor mature — 30代女性、ニュースキャスター風の落ち着いた声. How this speaker was made The voice identity for Spk4 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk4-10k.audiotext-to-speech10K<n<100K0 likes104 downloads5mo agoHugging Face06SynDataLab-JA-Refs /Irodori-Ja-Spk3-10k SynDataLab/Irodori-Ja-Spk3-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声. How this speaker was made The voice identity for Spk3 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k.audiotext-to-speech10K<n<100K0 likes95 downloads5mo agoHugging Face07hypaai /Hypa-Speech-10k A multilingual instruction-tuning dataset covering translation,transcription, and language detection. Dataset Card for Hypa-Speech-10k Dataset Summary Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets. The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Speech-10k.audioautomatic-speech-recognition10K<n<100K0 likes88 downloads3mo agoHugging Face08Vikhrmodels /Ficbook-Audio-Instruct-10K Ficbook Audio Instruct 10K Synthetic audio instruction dataset for training Russian audio-language models. Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks. Dataset Description This dataset was created for training and evaluating audio-language models on Russian fiction content. Each sample contains: Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model Text: Original text from ficbook stories Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.audioautomatic-speech-recognition1K<n<10K0 likes82 downloads9mo agoHugging Face09Duckyle /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Duckyle/meow-10k.audio10K<n<100K0 likes77 downloads4mo agoHugging Face10SynDataLab-JA-Refs /irodori-refs-10k Irodori TTS Reference Voices (10K) 10,000 synthetic Japanese reference voices generated with the Irodori-TTS-500M-v2-VoiceDesign model from voice-design captions (no reference audio — no_ref=True). Each row is one unique speaker. Columns column type description audio Audio(48kHz mono) reference waveform text string Japanese utterance with emoji prosody cues speaker_id string speaker_00001 … speaker_10000 Related datasets Clones… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/irodori-refs-10k.audiotext-to-speech10K<n<100K0 likes69 downloads5mo agoHugging Face11SynDataLab-JA-Refs /Irodori-Ja-Spk1-10k SynDataLab/Irodori-Ja-Spk1-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk1): 30s male, calm conversational — 30代男性、落ち着いた自然な会話調. How this speaker was made The voice identity for Spk1 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk1-10k.audiotext-to-speech10K<n<100K0 likes66 downloads5mo agoHugging Face12akuzdeuov /turkish_male_10kaudio10K<n<100K0 likes57 downloads10mo agoHugging Face13Kppwdfgu1 /Hypa-Speech-10k A multilingual instruction-tuning dataset covering translation,transcription, and language detection. Dataset Card for Hypa-Speech-10k Dataset Summary Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets. The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/Hypa-Speech-10k.audioautomatic-speech-recognition10K<n<100K0 likes55 downloads3mo agoHugging Face14SynDataLab-JA-Refs /irodori-refs-10k-v2 Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no_ref=True) using a richer caption space than v1: 8 axes (gender × age × pitch × tone × speed × distance × emotion × quality) with per-voice unique caption combinations, plus gender alternation, an incompatibility filter (no contradictory "whisper + speak loudly" combos), and a similarity-rejection window so consecutive voices stay distinct. Each ref's text… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/irodori-refs-10k-v2.audiotext-to-speech10K<n<100K0 likes54 downloads5mo agoHugging Face15akuzdeuov /turkish_female_10kaudio10K<n<100K0 likes51 downloads10mo agoHugging Face16assoni2002 /jailbreak_with_features_10kaudio10K<n<100K0 likes48 downloads1y agoHugging Face17SynDataLab-JA-Refs /Irodori-Ja-Spk2-10k SynDataLab/Irodori-Ja-Spk2-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk2): 30s female, narrator-style natural — 30代女性、ナレーター風の自然な声. How this speaker was made The voice identity for Spk2 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk2-10k.audiotext-to-speech10K<n<100K0 likes48 downloads5mo agoHugging Face18lokinfey /vibecoding-chinese-audio-book-tts-10kaudio10K<n<100K0 likes45 downloads10mo agoHugging Face19Rakancorle1 /hans-10k Hans-10K · DPO recipe for the audio-visual Clever Hans DPO training data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans 🐎 — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-10K is the 10,383-sample best-recipe preference-pair dataset that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.audioaudio-classification10K<n<100K0 likes45 downloads4mo agoHugging Face20lonesamurai /emilia_clean_10k EMILIA Clean 10k A filtered subset of the amphion/Emilia-Dataset (English split), designed for single-speaker TTS training. Dataset Statistics Total clips: 10,000 Speakers: 200 (single-speaker English) Train / Val split: 8,000 / 2,000 Duration per clip: 3–10 seconds Sample rate: 24 kHz (mono) Language: English (EN) Filtering Pipeline Candidate selection — Filtered EMILIA EN clips for duration (3–10s) and DNSMOS quality (≥3.2). Selected top 400 speakers with… See the full description on the dataset page: https://huggingface.co/datasets/lonesamurai/emilia_clean_10k.audio10K<n<100K1 likes42 downloads5mo agoHugging Face21shoron08 /irodori-refs-10k Irodori TTS Reference Voices (10K) 10,000 synthetic Japanese reference voices generated with the Irodori-TTS-500M-v2-VoiceDesign model from voice-design captions (no reference audio — no_ref=True). Each row is one unique speaker. Columns column type description audio Audio(48kHz mono) reference waveform text string Japanese utterance with emoji prosody cues speaker_id string speaker_00001 … speaker_10000 Related datasets Clones… See the full description on the dataset page: https://huggingface.co/datasets/shoron08/irodori-refs-10k.audiotext-to-speech10K<n<100K0 likes36 downloads4mo agoHugging Face22RidheshBhati /filipino-tts-10k-finalaudio10K<n<100K0 likes34 downloads8mo agoHugging Face23glenn2 /mls_eng_10k_train_part_1audio100K<n<1M0 likes32 downloads1y agoHugging Face24amine-maazizi /vc-detection-10kaudioaudio-classification10K<n<100K0 likes27 downloads6mo agoHugging Face25chuyangchenn /a-tre-10k A-TRE-10k Audio Tree Reconstruction Error benchmark — 10,000 synthetic audio scenes for evaluating whether audio encoders represent multi-source scenes compositionally. Companion dataset to the ICASSP 2026 paper Evaluating Compositional Structure in Audio Representations. See also the zero-shot benchmark chuyangchenn/a-coat-2k. Quick start from datasets import load_dataset ds = load_dataset("chuyangchenn/a-tre-10k", split="train") # or "val", "test" ex = ds[0]… See the full description on the dataset page: https://huggingface.co/datasets/chuyangchenn/a-tre-10k.audioaudio-classification10K<n<100K0 likes23 downloads5mo agoHugging Face26ckadirt /auramix_10kl auramix_10kl AuraMix is a small curated audio reconstruction/evaluation mix generated from multiple Hugging Face audio sources. Dataset summary Repo: ckadirt/auramix_10kl Clips: 10000 WAV files: 10000 Approx local size: 49.29 GB Sample rate: 44100 Clip duration: 60.0 seconds Mono: True Sources fma_full: 6000 clips from benjamin-paine/free-music-archive-full fma_commercial_full: 4000 clips from benjamin-paine/free-music-archive-commercial-16khz-full… See the full description on the dataset page: https://huggingface.co/datasets/ckadirt/auramix_10kl.audio10K<n<100K0 likes20 downloads5mo agoHugging Face27ckadirt /auramix_10km auramix_10km AuraMix is a small curated audio reconstruction/evaluation mix generated from multiple Hugging Face audio sources. Dataset summary Repo: ckadirt/auramix_10km Clips: 8441 WAV files: 8441 Approx local size: 20.81 GB Sample rate: 44100 Clip duration: 30.0 seconds Mono: True Sources fma_small: 2600 clips from benjamin-paine/free-music-archive-small fma_medium: 2600 clips from benjamin-paine/free-music-archive-medium fma_commercial_full: 2200 clips… See the full description on the dataset page: https://huggingface.co/datasets/ckadirt/auramix_10km.audio1K<n<10K0 likes19 downloads5mo agoHugging Face28Dramazy /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Dramazy/meow-10k.audio10K<n<100K0 likes19 downloads4mo agoHugging Face29AhunInteligence /ft_read_10kaudio10K<n<100K0 likes12 downloads7mo agoHugging Face30GrigoriiA /libritts_r_dataset_10kaudio10K<n<100K0 likes11 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.