datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2a-daddy
t2a-daddy
Male-voice ASMR corpus for the text2asmr project.
Previously published as aoxo/audios3.
Companion repos: aoxo/t2a-mommy (female voice),
aoxo/t2a-audios-v1 (the original v1 corpus).
Layout
path
what
<creator>/<title>.m4a
source audio, 48 kHz AAC, one folder per creator
<creator>/<title>.json
word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable)
labels/qwen3omni.jsonl
non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-daddy.atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.Urbansound8K_t2a
Dataset Card for "Urbansound8K_t2a"
More Information needed
MACS_t2a
Dataset Card for "MACS_t2a"
More Information needed
gigaspeech_t2aspoken-squad-t2aatlas-24-frozen-prefix-potential-shaping
ATLAS report 24: frozen-prefix potential shaping
1. Question and links
Read this first. This data root holds the first attempt of report 24 on the campaign's old harness (verl 0.7.1): the shaped training is complete and the unshaped training stopped at step 20 with a known problem (the subsection at the end of this section). The question was rerun on the runtime of report 25 with both trainings at 40 steps; that rerun's trajectories, exports, checkpoints and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-frozen-prefix-potential-shaping.T2AV_TemporalConstraintTalker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/: directories… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Talker-T2AV-Data.T2AV-Compass
Dataset Card for T2AV-Compass
Dataset Details
Dataset Description
T2AV-Compass is a unified benchmark for evaluating Text-to-Audio-Video (T2AV) generation, targeting not only unimodal quality (video/audio) but also cross-modal alignment & synchronization, complex instruction following, and perceptual realism grounded in physical/common-sense constraints.
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/T2AV-Compass.atlas-16-verifier-permission-prompt-ablation
ATLAS report 16: does the orchestrator's "cannot solve" clause suppress candidate verification?
Complete raw products of the ATLAS rl-training report 16 experiment
(GitHub issue #36). Two system-prompt arms of the same model over the
same 78 fixed states, greedy decoding, one shared vLLM server.
What the experiment did
The ATLAS orchestrator's frozen system prompt contains the clause
You cannot solve the problem yourself; you decide when to explore
further and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-16-verifier-permission-prompt-ablation.t2adata3atlas-15-selector-capacity-vs-training
ATLAS report 15 -- selector capacity vs training (raw products)
Every raw product of ATLAS rl-training report 15, uploaded for the raw-trace
audit asked for in t2ance/ATLAS issue #13.
Nothing is filtered or hand-picked: all 78 states are here for all five model
rows, both surfaces, both modes, both rollouts.
The experiment in one paragraph
39 LiveCodeBench questions whose 8-candidate Sonnet cache holds at least one
correct and one incorrect candidate that carries… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-15-selector-capacity-vs-training.sounddescs_t2a
SoundDescsT2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Natural language description for different audio sources from the BBC Sound Effects webpage.
Task category
Any2AnyRetrieval (text-to-audio)
Domains
Encyclopaedic, Written
Reference
IEEE Transactions on Multimedia
Source datasets:
mteb/sounddescs_t2a
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/sounddescs_t2a.audiocaps_t2a
AudioCapsT2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Natural language description for any kind of audio in the wild.
Task category
t2a
Domains
Encyclopaedic, Written
Reference
https://audiocaps.github.io/
Source datasets:
mteb/audiocaps_t2a
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("AudioCapsT2ARetrieval")
evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/audiocaps_t2a.atlas-23-prefix-curves-on-a-larger-selector
ATLAS report 23: prefix curves on a larger selector
1. Question and links
Read this first. GPQA ran in full on both surfaces (198 questions, prefix lengths k = 1 to 8, 1584 states each). LiveCodeBench was started and stopped by the user at 455 and 531 of its 1400 states per surface and is not read: every table and the report read GPQA only. A reader who cannot fetch files from the Hub finds this whole data root mirrored in the private GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-23-prefix-curves-on-a-larger-selector.mat-01-luna-teacher-search
01 Luna teacher search
Does a much larger black-box model, reached through a local proxy with no code change to the campaign's own tree-search harness, sample solution families other than the Qwen3.6-27B student, and how strong are its scores against the tasks' medal lines and against the student's recorded archives, and is its trajectory archive worth distilling into the student? This repository is the data root of that question: everything its ten standalone harness.run trees… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/mat-01-luna-teacher-search.clotho_t2a_v2
ClothoT2ARetrieval.v2
An MTEB dataset
Massive Text Embedding Benchmark
An audio captioning dataset containing audio clips from the Freesound platform and their corresponding captions. Version 2 removes empty-string queries. For more information see #5062
Task category
Any2AnyRetrieval (text-to-audio)
Domains
Encyclopaedic, Written
Reference
Clotho: An Audio Captioning Dataset
Source datasets:
mteb/Clotho
mteb/Clotho
How to evaluate on this task… See the full description on the dataset page: https://huggingface.co/datasets/lxercode/clotho_t2a_v2.dancegrpo-t2av
MiniMax H3 FL2VA First-Frame Dataset
27,815 FLUX-generated reference images paired with text prompts, built as
first-frame (image) conditions for MiniMax H3 FL2VA (text+image to
audio-video) RL training.
Dataset recipe
Prompts (prompts.txt, 27,815 lines): English video captions from
ConsisID-preview-Data,
as filtered and released by DanceGRPO
(assets/consist-id.txt).
Images (images/{index:06d}.jpg): each prompt rendered offline with
FLUX.1-dev on 8 GPUs — 400x640… See the full description on the dataset page: https://huggingface.co/datasets/zyfenghit/dancegrpo-t2av.mle-playbooksatlas-29-supergpqa-records-as-a-candidate-pool
29. Is SuperGPQA-Records a candidate pool an orchestrator can learn from?
1. Question and links
Does a candidate pool built from SuperGPQA-Records (128 hard questions, eight candidates each from the eight strongest single-shot models of the records) leave room for an untrained orchestrator, Qwen3.6-27B or Qwen3.5-9B, to learn: how does each model's greedy accuracy move from no candidate to eight against the pool's coverage? Report source:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-29-supergpqa-records-as-a-candidate-pool.atlas-24-validation-trajectories
ATLAS report 24: validation trajectories of both arms
Every validation trajectory that the two training runs of report 24 (frozen-prefix potential shaping) dumped, as one ordinary Hugging Face dataset: two splits, one per arm, one row per (validation step, question).
shaped: the shaped arm (W&B run 7mabv2p1), validations at steps 0, 5, 10, 15, 20, 25, 30, 35, 40 (3357 rows).
baseline: the baseline arm (W&B run 3fez6lk9), validations at steps 0, 5, 10, 15, 20 (1865 rows). The… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-validation-trajectories.VGGSound-T2AVThis is the VGGSound dataset (annotated with video and audio prompts) for paper "Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation"
This repo only contains the annotated train and evaluation metadata, please download the video files from Loie/VGGSound.
arXiv: https://arxiv.org/abs/2512.02457
Project: https://jianzongwu.github.io/projects/does-hearing-help-seeing/
Code: https://github.com/jianzongwu/Does-Hearing-Help-Seeing
jl_corpus_t2aTalker-T2AV-Data_trainer
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/prakhar-adaf/Talker-T2AV-Data_trainer.SongDescriber-T2ATalker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/Prakhar-kumar/Talker-T2AV-Data.t2adata2MusicCaps_t2a
Dataset Card for "MusicCaps_t2a"
More Information needed
LPMusicCapsMTT_t2a
LPMusicCapsMTTT2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
LLM-generated pseudo captions for 10-second music clips from the MagnaTagATune dataset. Captions were produced by prompting a large language model with the human-annotated tags of each clip, giving four differently-styled captions per clip. Complements MusicCaps, whose captions are human-written and whose audio comes from AudioSet.
Task category
Any2AnyRetrieval (text-to-audio)
Domains
Music… See the full description on the dataset page: https://huggingface.co/datasets/hubxrt/LPMusicCapsMTT_t2a.
