datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.open-vi-dialog-synthetic-100h
OpenDialog Vietnamese Synthetic Dialogue 100h
Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments.
12,000 chunks
30 seconds per chunk
100.0 hours total
Each item contains S1/S2 speaker labels, turn timings, target text,
relationship, pronouns, environment, topic, mood, and source reference IDs.
Audio renderer: vLLM-Omni VoxCPM2
Audio format: mono WAV, 48 kHz, 30 seconds per chunk
This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.rgad-crosslingual-tts-10h
RGAD Cross-Lingual TTS 10h
This is a 10-hour cross-lingual TTS dataset for prompt-conditioned Chinese TTS fine-tuning.
Format
The dataset contains:
train.jsonl
dev.jsonl
metadata.csv
audio/prompts/*.wav
audio/targets/*.wav
Each JSONL row has this format:
{"id":"sample_000001","prompt_wav":"audio/prompts/sample_000001.wav","target_wav":"audio/targets/sample_000001.wav","text":"中文目标文本。","prompt_language":"en-US","target_language":"zh-CN"… See the full description on the dataset page: https://huggingface.co/datasets/isabeth/rgad-crosslingual-tts-10h.meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Duckyle/meow-10k.aishell1mix-ver2-n100-per-mix
AISHELL-1 Mix ver2 — 100 clips per mix
This is AISHELL-1 Mix ver2, not ver1. 8 kHz mono test subset: 100 mixtures per speaker count (N=1\ldots5) (50 mix_clean + 50 mix_both each) → 500 clips.
Derived from the local aishell1mix_ver2 test SCPs (data/scp/scp_aishell1mix_ver2). Includes mixture + oracle speaker stems and transcripts.
Split
Count
1mix / 2mix / 3mix / 4mix / 5mix
100 each
clean / both
250 each
Files
manifests/test.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/aishell1mix-ver2-n100-per-mix.hans-10k
Hans-10K · DPO recipe for the audio-visual Clever Hans
DPO training data accompanying the paper
When Vision Speaks for Sound.
Like the original Clever Hans 🐎 —
the horse that looked like he could do arithmetic but was actually reading
his trainer's body language — video-capable MLLMs often look like they
can hear: they answer audio questions by reading visual cues and never
verifying the audio stream.
Hans-10K is the 10,383-sample best-recipe preference-pair dataset
that cures this… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-10k.CaptionStew10M-Qwen3Omnimeow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Dramazy/meow-10k.itts-controllability
CommonVoice ITTS Controllability Validation Set
Ground-truth-labelled English speech clips drawn from
fixie-ai/common_voice_17_0
(Common Voice 17.0, CC0), used to validate audio-LM judges for
instructable-TTS (ITTS) controllability scoring on age, gender, and
native accent.
Each clip carries a human-provided attribute label from Common Voice, so a judge's
perceived-attribute accuracy can be measured against real ground truth (blind
perceived-accuracy protocol).… See the full description on the dataset page: https://huggingface.co/datasets/Snooow1029/itts-controllability.
