datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.atlas-24-frozen-prefix-potential-shaping
ATLAS report 24: frozen-prefix potential shaping
1. Question and links
Read this first. This data root holds the first attempt of report 24 on the campaign's old harness (verl 0.7.1): the shaped training is complete and the unshaped training stopped at step 20 with a known problem (the subsection at the end of this section). The question was rerun on the runtime of report 25 with both trainings at 40 steps; that rerun's trajectories, exports, checkpoints and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-frozen-prefix-potential-shaping.Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/: directories… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Talker-T2AV-Data.atlas-23-prefix-curves-on-a-larger-selector
ATLAS report 23: prefix curves on a larger selector
1. Question and links
Read this first. GPQA ran in full on both surfaces (198 questions, prefix lengths k = 1 to 8, 1584 states each). LiveCodeBench was started and stopped by the user at 455 and 531 of its 1400 states per surface and is not read: every table and the report read GPQA only. A reader who cannot fetch files from the Hub finds this whole data root mirrored in the private GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-23-prefix-curves-on-a-larger-selector.VGGSound-T2AVThis is the VGGSound dataset (annotated with video and audio prompts) for paper "Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation"
This repo only contains the annotated train and evaluation metadata, please download the video files from Loie/VGGSound.
arXiv: https://arxiv.org/abs/2512.02457
Project: https://jianzongwu.github.io/projects/does-hearing-help-seeing/
Code: https://github.com/jianzongwu/Does-Hearing-Help-Seeing
Talker-T2AV-Data_trainer
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/prakhar-adaf/Talker-T2AV-Data_trainer.Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/Prakhar-kumar/Talker-T2AV-Data.t2adata2Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/dsenflam/Talker-T2AV-Data.
