datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.bangla-10k
Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh
Bangla-10K is a 10,070.8-hour Bengali speech corpus with
567,323 recordings from India and Bangladesh. It combines scripted
single-speaker read speech with natural multi-speaker conversations for
Bengali automatic speech recognition (ASR). The paper rounds the corpus scale
to 10,000 hours.
The corpus and its ASR evaluation are described in the anonymous manuscript… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.vls-10k
VLS 10K
9,987 MS-COCO images paired with a long written description, a one-sentence
spoken summary of that description, the spoken audio, and that audio pre-encoded
by two neural codecs. Images and audio are embedded in the parquet, so the
viewer renders them and one call opens the set:
from datasets import load_dataset
ds = load_dataset("seonglae/vls-10k", split="train")
ds[0]["image"] # PIL image
ds[0]["audio"] # decoded waveform
ds[0]["sst"] # the sentence that was… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/vls-10k.Irodori-Ja-Spk4-10k
SynDataLab/Irodori-Ja-Spk4-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk4): 30s female, news-anchor mature — 30代女性、ニュースキャスター風の落ち着いた声.
How this speaker was made
The voice identity for Spk4 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk4-10k.Irodori-Ja-Spk3-10k
SynDataLab/Irodori-Ja-Spk3-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声.
How this speaker was made
The voice identity for Spk3 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k.Hypa-Speech-10k
A multilingual instruction-tuning dataset covering translation,transcription, and language detection.
Dataset Card for Hypa-Speech-10k
Dataset Summary
Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets.
The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Speech-10k.Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Duckyle/meow-10k.irodori-refs-10k
Irodori TTS Reference Voices (10K)
10,000 synthetic Japanese reference voices generated with the
Irodori-TTS-500M-v2-VoiceDesign
model from voice-design captions (no reference audio — no_ref=True).
Each row is one unique speaker.
Columns
column
type
description
audio
Audio(48kHz mono)
reference waveform
text
string
Japanese utterance with emoji prosody cues
speaker_id
string
speaker_00001 … speaker_10000
Related datasets
Clones… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/irodori-refs-10k.Irodori-Ja-Spk1-10k
SynDataLab/Irodori-Ja-Spk1-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk1): 30s male, calm conversational — 30代男性、落ち着いた自然な会話調.
How this speaker was made
The voice identity for Spk1 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk1-10k.turkish_male_10kHypa-Speech-10k
A multilingual instruction-tuning dataset covering translation,transcription, and language detection.
Dataset Card for Hypa-Speech-10k
Dataset Summary
Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets.
The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/Hypa-Speech-10k.irodori-refs-10k-v2
Irodori TTS Reference Voices v2 (10K)
10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign
(no_ref=True) using a richer caption space than v1: 8 axes (gender × age ×
pitch × tone × speed × distance × emotion × quality) with per-voice unique
caption combinations, plus gender alternation, an incompatibility filter (no
contradictory "whisper + speak loudly" combos), and a similarity-rejection
window so consecutive voices stay distinct.
Each ref's text… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/irodori-refs-10k-v2.turkish_female_10kjailbreak_with_features_10kIrodori-Ja-Spk2-10k
SynDataLab/Irodori-Ja-Spk2-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk2): 30s female, narrator-style natural — 30代女性、ナレーター風の自然な声.
How this speaker was made
The voice identity for Spk2 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk2-10k.vibecoding-chinese-audio-book-tts-10khans-10k
Hans-10K · DPO recipe for the audio-visual Clever Hans
DPO training data accompanying the paper
When Vision Speaks for Sound.
Like the original Clever Hans 🐎 —
the horse that looked like he could do arithmetic but was actually reading
his trainer's body language — video-capable MLLMs often look like they
can hear: they answer audio questions by reading visual cues and never
verifying the audio stream.
Hans-10K is the 10,383-sample best-recipe preference-pair dataset
that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.emilia_clean_10k
EMILIA Clean 10k
A filtered subset of the amphion/Emilia-Dataset (English split), designed for single-speaker TTS training.
Dataset Statistics
Total clips: 10,000
Speakers: 200 (single-speaker English)
Train / Val split: 8,000 / 2,000
Duration per clip: 3–10 seconds
Sample rate: 24 kHz (mono)
Language: English (EN)
Filtering Pipeline
Candidate selection — Filtered EMILIA EN clips for duration (3–10s) and DNSMOS quality (≥3.2). Selected top 400 speakers with… See the full description on the dataset page: https://huggingface.co/datasets/lonesamurai/emilia_clean_10k.irodori-refs-10k
Irodori TTS Reference Voices (10K)
10,000 synthetic Japanese reference voices generated with the
Irodori-TTS-500M-v2-VoiceDesign
model from voice-design captions (no reference audio — no_ref=True).
Each row is one unique speaker.
Columns
column
type
description
audio
Audio(48kHz mono)
reference waveform
text
string
Japanese utterance with emoji prosody cues
speaker_id
string
speaker_00001 … speaker_10000
Related datasets
Clones… See the full description on the dataset page: https://huggingface.co/datasets/shoron08/irodori-refs-10k.filipino-tts-10k-finalmls_eng_10k_train_part_1vc-detection-10ka-tre-10k
A-TRE-10k
Audio Tree Reconstruction Error benchmark — 10,000 synthetic audio scenes
for evaluating whether audio encoders represent multi-source scenes compositionally.
Companion dataset to the ICASSP 2026 paper Evaluating Compositional Structure in Audio
Representations. See also the
zero-shot benchmark chuyangchenn/a-coat-2k.
Quick start
from datasets import load_dataset
ds = load_dataset("chuyangchenn/a-tre-10k", split="train") # or "val", "test"
ex = ds[0]… See the full description on the dataset page: https://huggingface.co/datasets/chuyangchenn/a-tre-10k.auramix_10kl
auramix_10kl
AuraMix is a small curated audio reconstruction/evaluation mix generated from multiple Hugging Face audio sources.
Dataset summary
Repo: ckadirt/auramix_10kl
Clips: 10000
WAV files: 10000
Approx local size: 49.29 GB
Sample rate: 44100
Clip duration: 60.0 seconds
Mono: True
Sources
fma_full: 6000 clips from benjamin-paine/free-music-archive-full
fma_commercial_full: 4000 clips from benjamin-paine/free-music-archive-commercial-16khz-full… See the full description on the dataset page: https://huggingface.co/datasets/ckadirt/auramix_10kl.auramix_10km
auramix_10km
AuraMix is a small curated audio reconstruction/evaluation mix generated from multiple Hugging Face audio sources.
Dataset summary
Repo: ckadirt/auramix_10km
Clips: 8441
WAV files: 8441
Approx local size: 20.81 GB
Sample rate: 44100
Clip duration: 30.0 seconds
Mono: True
Sources
fma_small: 2600 clips from benjamin-paine/free-music-archive-small
fma_medium: 2600 clips from benjamin-paine/free-music-archive-medium
fma_commercial_full: 2200 clips… See the full description on the dataset page: https://huggingface.co/datasets/ckadirt/auramix_10km.meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Dramazy/meow-10k.ft_read_10klibritts_r_dataset_10k
