datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.prostate128_t2_anatomy_nnUNet_3d_fullres_20_epocht2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.FoR-T2I
Can Text-to-Image Models Draw from the Right Frame of Reference?
Left is not always image-left. If a person is facing the viewer, their left hand appears on the right side of the image. A text-to-image model can therefore render a plausible scene with the requested objects while still drawing the spatial relation from the wrong perspective.
FoR-T2I is the official benchmark release for Can Text-to-Image Models Draw from the Right Frame of Reference?. It… See the full description on the dataset page: https://huggingface.co/datasets/ernie-research/FoR-T2I.T2I-CoReBench
Easier Painting Than Thinking: Can Text-to-Image Models
Set the Stage, but Not Direct the Play?
Ouxiang Li1*, Yuan Wang1, Xinting Hu†, Huijuan Huang2‡, Rui Chen2, Jiarong Ou2,
Xin Tao2†, Pengfei Wan2, Xiaojuan Qi3, Fuli Feng1
1University of Science and Technology of China, 2Kling Team, Kuaishou Technology, 3The University of Hong Kong
*Work done during internship… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/grasson/t2-ragbench.DIM-T2I
[ICLR 2026] Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
Ziyun Zeng, David Junhao Zhang, Wei Li,
and Mike Zheng Shou
📰 News
[2026-05-12] The DIM project page is available.
[2026-01-26] 🎉 DIM is accepted to ICLR 2026!
[2025-10-08] 🚀 Released the DIM-Edit dataset and the DIM-4.6B-T2I / DIM-4.6B-Edit models.
[2025-09-02] 📝 The DIM paper is released on arXiv.
🌟 Highlights
🧠… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/DIM-T2I.Sera-4.6-Lite-T2This dataset contains 36083 trajectories. A 25000 subset was used to train SERA-32B. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.6 as teacher.
Schema:
messages: Generated trajectory
instance_id: ID of trajectory
rollout_patch: Created patch to the codebase from the current trajectory
func_name: Name of function sampled from codebase to start the pipeline
func_path: File path to the sampled function
problem_statement: Problem statement provided to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.6-Lite-T2.Sera-4.5A-Full-T2This dataset contains 66337 trajectories. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes three SVG runs per function. Sera-4.5-Lite-T2 is a subset of this dataset and was used to train SERA-32B-GA.
Schema:
messages: Generated trajectory
instance_id: ID of trajectory
rollout_patch: Created patch to the codebase from the current trajectory
func_name: Name of function sampled from codebase to start the pipeline
func_path:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Full-T2.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/tomsummerfield/t2-ragbench.sorrel-T2-qwen3-8b-base-seed0-documentssorrel-T2-qwen3-4b-base-seed0-documents2026-08-22-ruleform-ablated2-t2-9284-synthdoc-676
Rewrite-then-delete, two ablation passes, responsiveness-filtered (Table2 9,284 + difficult-advice 676)
field
value
experiment
Ablation arm built in two passes over BOTH halves of every difficult-advice row. Pass 1 rewrites the reasoning and the answer to drop four deliberative moves: engaging the tempting option, drawing an analytic distinction, enumerating outcome branches, and offering an alternative route. Pass 2 then labels every remaining unit and DELETES the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-22-ruleform-ablated2-t2-9284-synthdoc-676.2026-08-28-t2-9284-attack716-train
Attack-variant robustness mixture (9,284 + 716)
field
value
experiment
Training mixture for the attack-variant robustness arm. 179 difficult-advice scenarios x four user-prompt framings that all press for the same norm-violating shortcut (original, hidden, incremental, authority); the assistant reply is held constant across a scenario's four framings (the correct refusal). Plus the same 9,284 Table2 rows. Trains the model to hold its line when the ask is reframed.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/2026-08-28-t2-9284-attack716-train.azm-archive-20260909-t2v-needlabel-videos
t2v_data_needlabel_videos.tar
Backup of an existing dataset archive, preserving its original bytes.
File: t2v_data_needlabel_videos.tar
Size: 4,675,983,360 bytes
SHA256: 8672488cef23b55e2a70d3279327a38ddbbd8e2809661600d7c707d59b9e1e18
Verify the downloaded archive with sha256sum -c SHA256SUMS.
Arena-T2I-Hard
Arena-T2I-Hard
A 310-prompt stress benchmark for evaluating faithfulness (prompt-following) of
text-to-image models, drawn from real, hard arena user requests — long, multi-entity
prompts with attributes, spatial relations, counts, and stylistic constraints. Each
prompt ships pre-decomposed into a dependency-aware DAG of yes/no questions; when
scoring an image, failing a parent question zeroes out its descendants. The benchmark
stays discriminative where DPG-Bench and DSG… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/Arena-T2I-Hard.Info_Wan_Video_2.2_T2V-A14B
Model Index by Creator
423748 Page
Model
Base Model
Full Model Page
Archive Link
wan2.2,t2v,low,zzzyixuan.
Wan Video 2.2 T2V-A14B
View
View
Version Links
Model
Version
Base Model
Version Link
wan2.2,t2v,low,zzzyixuan.
v1.0
Wan Video 2.2 T2V-A14B
View
Aaron_PP Page
Model
Base Model
Full Model Page
Archive Link
NSFW WAN 2.2 T2V Bunny girl, red patent leather tights, black high stockings, red high heels
Wan Video 2.2… See the full description on the dataset page: https://huggingface.co/datasets/ApacheOne/Info_Wan_Video_2.2_T2V-A14B.Sera-4.5A-Lite-T2This dataset contains 35615 trajectories. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes one SVG run per function. 16000 samples from the dataset were used to train SERA-32B-GA. Sera-4.5-Full-T2 is a superset of this dataset with three SVG runs per function.
Schema:
messages: Generated trajectory
instance_id: ID of trajectory
rollout_patch: Created patch to the codebase from the current trajectory
func_name: Name of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Lite-T2.t2mle-playbooksAI4free__t2-details
Dataset Card for Evaluation run of AI4free/t2
Dataset automatically created during the evaluation run of model AI4free/t2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration "results"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI4free__t2-details.dancegrpo-t2av
MiniMax H3 FL2VA First-Frame Dataset
27,815 FLUX-generated reference images paired with text prompts, built as
first-frame (image) conditions for MiniMax H3 FL2VA (text+image to
audio-video) RL training.
Dataset recipe
Prompts (prompts.txt, 27,815 lines): English video captions from
ConsisID-preview-Data,
as filtered and released by DanceGRPO
(assets/consist-id.txt).
Images (images/{index:06d}.jpg): each prompt rendered offline with
FLUX.1-dev on 8 GPUs — 400x640… See the full description on the dataset page: https://huggingface.co/datasets/zyfenghit/dancegrpo-t2av.2026-08-28-t2-9284-advice716-train
Advice-framed difficult-advice mixture (9,284 + 716)
field
value
experiment
Training mixture for the advice-framed arm. The 716-row difficult-advice corpus behind qwen3.6-27b-lora-t2-9284-synthdoc-716-dynbatch-r64, lightly rewritten so the user asks whether to take the tempting shortcut and the assistant purely advises (framing only; wording/stance preserved). Plus the same 9,284 Table2 rows. A downstream difference against the synthdoc-716 arm isolates the… See the full description on the dataset page: https://huggingface.co/datasets/matboz/2026-08-28-t2-9284-advice716-train.sorrel-T2-qwen3.8-27b-sorrel-sdf-seed0-documentssorrel-T2-gemma-4-12b-seed0-documentsSera-4.5A-Django-T2This dataset contains 21900 trajectories. Data was generated from the second rollout of SVG on 6 Django commits using GLM-4.5-Air as teacher and includes one SVG runs per function.
We only run verification at 0.5 recall for specialization rollouts (second rollout).
Schema:
messages: Generated trajectory
instance_id: ID of trajectory
rollout_patch: Created patch to the codebase
func_name: Name of function sampled from codebase to start the pipeline
func_path: File path to the sampled function… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Django-T2.2026-08-31-cot-only-supervision-t2-9284-synthdoc-716
CoT-only supervision mixture (Table2 9,284 + difficult-advice 716)
field
value
experiment
Arm: train the 716 difficult-advice rows on their REASONING ONLY — each row is truncated at its </think> close, so the visible answer leaves both the loss and the forward pass — while the 9,284 Table2 rows train exactly as in the control. Tests whether the difficult-advice effect on agentic misalignment is carried by the reasoning or by the answer.
date_generated
2026-08-31… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-cot-only-supervision-t2-9284-synthdoc-716.T2I-Eval-BenchThis is the human-annotated benchmark dataset for paper Automatic Evaluation for Text-to-Image Generation: Fine-grained Framework,
Distilled Evaluation Model and Meta-Evaluation Benchmark
NOTE: Please check out our github repository for more detailed usage.
FlexiSLM-Data-5M-t2t
FlexiSLM-Data — Text-to-Text Part (5M)
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in
WebDataset format.
Related data releases
FlexiSLM/FlexiSLM-Data-5M-t2t (this repo) provides… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-5M-t2t.data-juicer-t2v-optimal-data-pool
Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development
Project description
The emergence of large-scale multi-modal generative models has drastically advanced artificial intelligence, introducing unprecedented levels of performance and functionality.
However, optimizing these models remains challenging due to historically isolated paths of model-centric and data-centric developments, leading to suboptimal outcomes and inefficient resource… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/data-juicer-t2v-optimal-data-pool.
