datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.atlas-24-frozen-prefix-potential-shaping
ATLAS report 24: frozen-prefix potential shaping
1. Question and links
Read this first. This data root holds the first attempt of report 24 on the campaign's old harness (verl 0.7.1): the shaped training is complete and the unshaped training stopped at step 20 with a known problem (the subsection at the end of this section). The question was rerun on the runtime of report 25 with both trainings at 40 steps; that rerun's trajectories, exports, checkpoints and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-frozen-prefix-potential-shaping.Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/: directories… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Talker-T2AV-Data.Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.bike-sharing-tabular
Bike Sharing Demand - Hourly (Poisson)
A ready-to-use copy of the UCI Bike Sharing Dataset (hourly granularity,
17,379 × 17), accompanied by baseline metrics from an 8-architecture tabular
modelling pipeline for direct comparison.
Originally collected and published by Fanaee-T & Gama (2014). Source:
UCI ML Repository id 275.
At a glance
Field
Value
Rows
17,379 hourly observations
Time range
Jan 2011 - Dec 2012
Columns
17 (16 features + 1 target)… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/bike-sharing-tabular.atlas-23-prefix-curves-on-a-larger-selector
ATLAS report 23: prefix curves on a larger selector
1. Question and links
Read this first. GPQA ran in full on both surfaces (198 questions, prefix lengths k = 1 to 8, 1584 states each). LiveCodeBench was started and stopped by the user at 455 and 531 of its 1400 states per surface and is not read: every table and the report read GPQA only. A reader who cannot fetch files from the Hub finds this whole data root mirrored in the private GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-23-prefix-curves-on-a-larger-selector.T2T-Centromere-Regulatory
T2T Centromere Regulatory
Curated and released by Basepair | Follow updates on X: @BasepairSci.
Dataset Summary
The T2T Centromere Regulatory is the first comprehensive, base-pair resolution mapping of cryptic transcriptional switches and secondary structural elements across the newly sequenced Telomere-to-Telomere (T2T-CHM13 v2.0 / hs1) human centromeres.
For decades, centromeric alpha-satellite DNA (~100–200 Mb across human chromosomes) was considered… See the full description on the dataset page: https://huggingface.co/datasets/Basepair/T2T-Centromere-Regulatory.house-prices-tabular
House Prices - Tabular (with baseline metrics)
A curated, ready-to-use copy of the Kaggle House Prices: Advanced
Regression Techniques training set (1,460 × 81), accompanied by baseline
metrics from an 8-architecture tabular modelling pipeline so newcomers
have a reference point to compare against.
This is the same data as Kaggle's train.csv, sourced from
OpenML id 42165 (canonical mirror).
At a glance
Field
Value
Rows
1,460
Columns
81 (80 features + 1… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/house-prices-tabular.hc01-t2d-sample
HC01 — Synthetic Type 2 Diabetes Patient Dataset (Evaluation Sample)
Publisher: XpertSystems.ai
SKU: HC01 (sample)
Version: 1.0.0
License: CC BY-NC 4.0 — non-commercial evaluation and research use only. Commercial use, redistribution, or derivative data products require a commercial license.
Full product: Contact pradeep@xpertsystems.ai
What this is
A 500-patient evaluation slice of the XpertSystems HC01 synthetic Type 2 Diabetes dataset, released for technical… See the full description on the dataset page: https://huggingface.co/datasets/xpertsystems/hc01-t2d-sample.VGGSound-T2AVThis is the VGGSound dataset (annotated with video and audio prompts) for paper "Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation"
This repo only contains the annotated train and evaluation metadata, please download the video files from Loie/VGGSound.
arXiv: https://arxiv.org/abs/2512.02457
Project: https://jianzongwu.github.io/projects/does-hearing-help-seeing/
Code: https://github.com/jianzongwu/Does-Hearing-Help-Seeing
Talker-T2AV-Data_trainer
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/prakhar-adaf/Talker-T2AV-Data_trainer.t2adata2Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/Prakhar-kumar/Talker-T2AV-Data.anonymization-before-after
Anonymization Before/After
A small paired tabular dataset showing the same records before and after
a 10-step anonymization pipeline. Useful as a teaching fixture for privacy
courses, a benchmark for anonymization toolkits, and a sanity-check input
for red-team / membership-inference experiments.
Important: the PII in sample_raw.csv is entirely synthetic.
Names follow the pattern Person_001, emails are person_001@example.com,
phone numbers are 555-00XX, and "national IDs" are… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/anonymization-before-after.Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/dsenflam/Talker-T2AV-Data.T2P
T2P: Textile-to-Physics fabric parameters
T2P is a tabular dataset of 1,382 real fabrics with physical properties expressed as
CLO3D cloth-simulation parameters, paired with each fabric's fiber composition and
construction metadata.
The task: predict a fabric's simulation-ready physical parameters (bending / shear /
stretch stiffness, buckling, friction, weight, damping) from its composition and
construction descriptors — bridging material identity ("95% cotton, 5% elastane… See the full description on the dataset page: https://huggingface.co/datasets/image2garment/T2P.T2G-1k-Qwen2.5-3B
T2G
Overview
T2G is a synthetic data consisting of text-graph pairs designed to finetune LLMs on information extraction tasks, specifically text-to-graph conversion.
Dataset Structure
The dataset is organized into the following main components:
Train Set: 800 instances for training models.
Validation Set: 100 for validating model performance.
Test Set: 100 instances for final evaluation.
Data Fields
Each instance in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/ESITime/T2G-1k-Qwen2.5-3B.t2i-diversity-gender-neutral-captionsThis dataset contains different synthetic captions for our image samples.
We have selected the best-performing caption set from our experiments, the random-length captions. Then, we have used Gemma-2-9b-it and instructed it to remove different genders from the captions. We obtained three sets from the original set, namly (i) all genders neutralized, (ii) only female gender neutralized, and (iii) only male gender neutralized. To this end, we have removed all gender indicative words such as… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/t2i-diversity-gender-neutral-captions.T2G-Event-K-HopDPT-T2I_training_datadiffing-stats-SAEdiff_ftb-qwen3_1_7B-kansas_abortion-L14-s1-t200-k100-lr1e-04-x32t2jsonMultilingual_T2I_clean_llama2_templated_promptsT2G-1k-Llama3.2-3B
T2G
Overview
T2G is a synthetic data consisting of text-graph pairs designed to finetune LLMs on information extraction tasks, specifically text-to-graph conversion.
Dataset Structure
The dataset is organized into the following main components:
Train Set: 800 instances for training models.
Validation Set: 100 for validating model performance.
Test Set: 100 instances for final evaluation.
Data Fields
Each instance in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/ESITime/T2G-1k-Llama3.2-3B.T2G-Event-1k-Qwen2.5-3B-changed-formatT20WinnerPredictor
T20 Winner Predictor Dataset
This repository contains a T20 match winner prediction dataset.
Columns
Match_ID
Winner
Target Column
Winner
T2G-Event-1k-Qwen2.5-3Bdiffing-stats-SAEdiff_ftb-qwen3_1_7B-kansas_abortion-L14-k100-x4-lr1e-04-t200diffing-stats-SAEdiff_ftb-qwen3_1_7B-kansas_abortion-L14-k100-x32-lr1e-04-t200t2
