CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01G4KMU /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.documenttable-question-answering10K<n<100K17 likes5.3k downloads6mo agoHugging Face02osbm /prostate128_t2_anatomy_nnUNet_3d_fullres_20_epochtextn<1K0 likes590 downloads3y agoHugging Face03botay /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.documenttable-question-answering10K<n<100K0 likes546 downloads5mo agoHugging Face04ernie-research /FoR-T2I Can Text-to-Image Models Draw from the Right Frame of Reference? Left is not always image-left. If a person is facing the viewer, their left hand appears on the right side of the image. A text-to-image model can therefore render a plausible scene with the requested objects while still drawing the spatial relation from the wrong perspective. FoR-T2I is the official benchmark release for Can Text-to-Image Models Draw from the Right Frame of Reference?. It… See the full description on the dataset page: https://huggingface.co/datasets/ernie-research/FoR-T2I.texttext-to-image1K<n<10K6 likes431 downloads2mo agoHugging Face05lioooox /T2I-CoReBench Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play? Ouxiang Li1*, Yuan Wang1, Xinting Hu†, Huijuan Huang2‡, Rui Chen2, Jiarong Ou2, Xin Tao2†, Pengfei Wan2, Xiaojuan Qi3, Fuli Feng1 1University of Science and Technology of China, 2Kling Team, Kuaishou Technology, 3The University of Hong Kong *Work done during internship… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench.imagetext-to-image1K<n<10K2 likes327 downloads7mo agoHugging Face06grasson /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/grasson/t2-ragbench.documenttable-question-answering10K<n<100K0 likes216 downloads5mo agoHugging Face07stdKonjac /DIM-T2I [ICLR 2026] Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing Ziyun Zeng, David Junhao Zhang, Wei Li, and Mike Zheng Shou 📰 News [2026-05-12] The DIM project page is available. [2026-01-26] 🎉 DIM is accepted to ICLR 2026! [2025-10-08] 🚀 Released the DIM-Edit dataset and the DIM-4.6B-T2I / DIM-4.6B-Edit models. [2025-09-02] 📝 The DIM paper is released on arXiv. 🌟 Highlights 🧠… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/DIM-T2I.imagetext-to-image1M<n<10M2 likes203 downloads5mo agoHugging Face08allenai /Sera-4.6-Lite-T2This dataset contains 36083 trajectories. A 25000 subset was used to train SERA-32B. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.6 as teacher. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path: File path to the sampled function problem_statement: Problem statement provided to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.6-Lite-T2.text10K<n<100K11 likes178 downloads7mo agoHugging Face09allenai /Sera-4.5A-Full-T2This dataset contains 66337 trajectories. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes three SVG runs per function. Sera-4.5-Lite-T2 is a subset of this dataset and was used to train SERA-32B-GA. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Full-T2.text10K<n<100K3 likes153 downloads7mo agoHugging Face10tomsummerfield /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/tomsummerfield/t2-ragbench.documenttable-question-answering10K<n<100K0 likes122 downloads6mo agoHugging Face11joshycodes /sorrel-T2-qwen3-8b-base-seed0-documentstext10K<n<100K0 likes105 downloads8d agoHugging Face12joshycodes /sorrel-T2-qwen3-4b-base-seed0-documentstext10K<n<100K0 likes105 downloads8d agoHugging Face13dougalldeepmind /2026-08-22-ruleform-ablated2-t2-9284-synthdoc-676 Rewrite-then-delete, two ablation passes, responsiveness-filtered (Table2 9,284 + difficult-advice 676) field value experiment Ablation arm built in two passes over BOTH halves of every difficult-advice row. Pass 1 rewrites the reasoning and the answer to drop four deliberative moves: engaging the tempting option, drawing an analytic distinction, enumerating outcome branches, and offering an alternative route. Pass 2 then labels every remaining unit and DELETES the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-22-ruleform-ablated2-t2-9284-synthdoc-676.text1K<n<10K0 likes103 downloads26d agoHugging Face14matboz /2026-08-28-t2-9284-attack716-train Attack-variant robustness mixture (9,284 + 716) field value experiment Training mixture for the attack-variant robustness arm. 179 difficult-advice scenarios x four user-prompt framings that all press for the same norm-violating shortcut (original, hidden, incremental, authority); the assistant reply is held constant across a scenario's four framings (the correct refusal). Plus the same 9,284 Table2 rows. Trains the model to hold its line when the ask is reframed.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/2026-08-28-t2-9284-attack716-train.text10K<n<100K0 likes93 downloads29d agoHugging Face15qgfvadfuvads /azm-archive-20260909-t2v-needlabel-videos t2v_data_needlabel_videos.tar Backup of an existing dataset archive, preserving its original bytes. File: t2v_data_needlabel_videos.tar Size: 4,675,983,360 bytes SHA256: 8672488cef23b55e2a70d3279327a38ddbbd8e2809661600d7c707d59b9e1e18 Verify the downloaded archive with sha256sum -c SHA256SUMS. textn<1K0 likes91 downloads17d agoHugging Face16lmarena-ai /Arena-T2I-Hard Arena-T2I-Hard A 310-prompt stress benchmark for evaluating faithfulness (prompt-following) of text-to-image models, drawn from real, hard arena user requests — long, multi-entity prompts with attributes, spatial relations, counts, and stylistic constraints. Each prompt ships pre-decomposed into a dependency-aware DAG of yes/no questions; when scoring an image, failing a parent question zeroes out its descendants. The benchmark stays discriminative where DPG-Bench and DSG… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/Arena-T2I-Hard.texttext-to-imagen<1K1 likes88 downloads3mo agoHugging Face17ApacheOne /Info_Wan_Video_2.2_T2V-A14B Model Index by Creator 423748 Page Model Base Model Full Model Page Archive Link wan2.2,t2v,low,zzzyixuan. Wan Video 2.2 T2V-A14B View View Version Links Model Version Base Model Version Link wan2.2,t2v,low,zzzyixuan. v1.0 Wan Video 2.2 T2V-A14B View Aaron_PP Page Model Base Model Full Model Page Archive Link NSFW WAN 2.2 T2V Bunny girl, red patent leather tights, black high stockings, red high heels Wan Video 2.2… See the full description on the dataset page: https://huggingface.co/datasets/ApacheOne/Info_Wan_Video_2.2_T2V-A14B.textn<1K7 likes84 downloads1y agoHugging Face18allenai /Sera-4.5A-Lite-T2This dataset contains 35615 trajectories. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes one SVG run per function. 16000 samples from the dataset were used to train SERA-32B-GA. Sera-4.5-Full-T2 is a superset of this dataset with three SVG runs per function. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Lite-T2.text10K<n<100K4 likes84 downloads7mo agoHugging Face19BrandDead /t2textn<1K0 likes83 downloads2y agoHugging Face20t2ance /mle-playbookstextn<1K0 likes83 downloads6mo agoHugging Face21open-llm-leaderboard /AI4free__t2-detailsgated Dataset Card for Evaluation run of AI4free/t2 Dataset automatically created during the evaluation run of model AI4free/t2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration "results"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI4free__t2-details.tabular10K<n<100K0 likes74 downloads2y agoHugging Face22zyfenghit /dancegrpo-t2av MiniMax H3 FL2VA First-Frame Dataset 27,815 FLUX-generated reference images paired with text prompts, built as first-frame (image) conditions for MiniMax H3 FL2VA (text+image to audio-video) RL training. Dataset recipe Prompts (prompts.txt, 27,815 lines): English video captions from ConsisID-preview-Data, as filtered and released by DanceGRPO (assets/consist-id.txt). Images (images/{index:06d}.jpg): each prompt rendered offline with FLUX.1-dev on 8 GPUs — 400x640… See the full description on the dataset page: https://huggingface.co/datasets/zyfenghit/dancegrpo-t2av.texttext-to-video10K<n<100K1 likes71 downloads1mo agoHugging Face23matboz /2026-08-28-t2-9284-advice716-train Advice-framed difficult-advice mixture (9,284 + 716) field value experiment Training mixture for the advice-framed arm. The 716-row difficult-advice corpus behind qwen3.6-27b-lora-t2-9284-synthdoc-716-dynbatch-r64, lightly rewritten so the user asks whether to take the tempting shortcut and the assistant purely advises (framing only; wording/stance preserved). Plus the same 9,284 Table2 rows. A downstream difference against the synthdoc-716 arm isolates the… See the full description on the dataset page: https://huggingface.co/datasets/matboz/2026-08-28-t2-9284-advice716-train.text10K<n<100K0 likes69 downloads29d agoHugging Face24joshycodes /sorrel-T2-qwen3.8-27b-sorrel-sdf-seed0-documentstext100K<n<1M0 likes67 downloads7d agoHugging Face25joshycodes /sorrel-T2-gemma-4-12b-seed0-documentstext10K<n<100K0 likes64 downloads8d agoHugging Face26allenai /Sera-4.5A-Django-T2This dataset contains 21900 trajectories. Data was generated from the second rollout of SVG on 6 Django commits using GLM-4.5-Air as teacher and includes one SVG runs per function. We only run verification at 0.5 recall for specialization rollouts (second rollout). Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase func_name: Name of function sampled from codebase to start the pipeline func_path: File path to the sampled function… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Django-T2.text10K<n<100K2 likes58 downloads8mo agoHugging Face27dougalldeepmind /2026-08-31-cot-only-supervision-t2-9284-synthdoc-716 CoT-only supervision mixture (Table2 9,284 + difficult-advice 716) field value experiment Arm: train the 716 difficult-advice rows on their REASONING ONLY — each row is truncated at its </think> close, so the visible answer leaves both the loss and the forward pass — while the 9,284 Table2 rows train exactly as in the control. Tests whether the difficult-advice effect on agentic misalignment is carried by the reasoning or by the answer. date_generated 2026-08-31… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-cot-only-supervision-t2-9284-synthdoc-716.text10K<n<100K0 likes57 downloads26d agoHugging Face28DataHammer /T2I-Eval-BenchThis is the human-annotated benchmark dataset for paper Automatic Evaluation for Text-to-Image Generation: Fine-grained Framework, Distilled Evaluation Model and Meta-Evaluation Benchmark NOTE: Please check out our github repository for more detailed usage. textn<1K3 likes54 downloads2y agoHugging Face29FlexiSLM /FlexiSLM-Data-5M-t2t FlexiSLM-Data — Text-to-Text Part (5M) Paper: https://arxiv.org/abs/2606.31247 Demo page: https://flexislm.github.io/ Code: https://github.com/AmphionTeam/FlexiSLM FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset for training FlexiSLM, a spoken language model. This repository contains the paired prompt-and-response audio portion of the release in WebDataset format. Related data releases FlexiSLM/FlexiSLM-Data-5M-t2t (this repo) provides… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-5M-t2t.texttext-generation1M<n<10M1 likes54 downloads2mo agoHugging Face30datajuicer /data-juicer-t2v-optimal-data-pool Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development Project description The emergence of large-scale multi-modal generative models has drastically advanced artificial intelligence, introducing unprecedented levels of performance and functionality. However, optimizing these models remains challenging due to historically isolated paths of model-centric and data-centric developments, leading to suboptimal outcomes and inefficient resource… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/data-juicer-t2v-optimal-data-pool.texttext-to-videon<1K0 likes50 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.