datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gastric-X
Gastric-X
Multi-phase abdominal CT cohort paired with structured laboratory panels
and free-text radiology reports, in proficient medical English with
the original Simplified Chinese preserved alongside.
Changelog
2026-06-26
Added per-phase organ masks (<phase>_organ_mask.nii.gz) — CADS
multi-organ segmentation on each phase's CT grid (e.g. label 6 = stomach);
all 4897 phases.
Added per-phase gastric tumor masks (<phase>_tumor_mask.nii.gz,
binary) — a patient's… See the full description on the dataset page: https://huggingface.co/datasets/HaoChen2/Gastric-X.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.MAPBench-V2For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.Reason-RFT-CoT-Dataset
🤗 Reason-RFT CoT Dateset
The full dataset used in our project "Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning".
⭐️ Project │ 🌎 Github │ 🔥 Models │ 📑 ArXiv │ 💬 WeChat
🤖 RoboBrain: Aim to Explore ReasonRFT Paradigm to Enhance RoboBrain's Embodied Reasoning Capabilities.
♣️ Quick Start
Please refer to Dataset Preparation
🔥 Overview
Visual reasoning abilities play a crucial role in understanding complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/tanhuajie2001/Reason-RFT-CoT-Dataset.DocDownstream-2.0DocDownstream 2.0 is a collection of MP-DocVQA, DUDE, NewsVideoQA used in DocOwl2.
For MP-DocVQA and DUDE, the maximum number of pages of each sample is set to 20.
For NewsVideoQA, the maximum number of frames of each sample is set to 20.
gdpval_preference_rubricsterminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version.
This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes:
Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.UAVid-2020
UAVid: Aerial Semantic Segmentation Dataset
Unofficial redistribution of the UAVid dataset under the original CC BY-NC-SA 4.0 license.
Disclaimer
This repository is not an official release of the UAVid dataset.
The UAVid dataset was created by the original authors, who retain all copyright and intellectual property rights. This repository does not claim ownership of any images, annotations, or metadata.
This repository exists for two purposes:
To… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/UAVid-2020.wan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.ProcVQA-20M-annotations
ProcVQA-20M Annotations
Project Page |
arXiv |
Code |
Model |
Media
This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media.
Overview
This dataset is constructed from over 26 embodied datasets, comprising:
20M QA pairs for training
330K original trajectories
50M annotated frames from ~5,000 hours of manipulation data
200+ different tasks
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.Minecraft-Skins-20M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 19,973,928 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
id: A randomly generated UUID for each skin entry. These UUIDs are not linked to any external APIs or services (such as Mojang's player UUIDs) and serve solely as… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/Minecraft-Skins-20M.pexels-568k-internvl2
Dataset Card for pexels-568k-internvl2
Dataset Summary
This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed.
Languages
The text is in English, but occasionally text in images in other languages is transcribed.
Intended Usage
Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.ArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/Thunderbolt215215/ArtiMuse-10K.Omni-Edu
Omni-Edu — Core V6 SFT mixture
69,999 supervised instruction examples (~158M characters) covering K-12 subject
competence, curriculum grounding, diagnostic reasoning, pedagogical action and
general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image
referenced by the JSONL ships in this repository under images/.
This is the system-prompted assembly of the v6 core mixture: every row carries
an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.SPADES-RGBmedical-prescription-datasetVisDocAgentBench
VisDocAgentBench
VisDocAgentBench is a closed-corpus benchmark for visually rich document retrieval. It contains 120 natural-language queries over 2,375 rendered pages from 100 scientific documents. Queries are evenly divided among direct, one-bridge, and two-bridge evidence structures, with one answer page per query.
[Paper] [Code] [Project page]
Contents
benchmark/
├── queries.jsonl
├── evaluator_annotations.jsonl
└── topics.json
corpus/
├── documents.jsonl
├──… See the full description on the dataset page: https://huggingface.co/datasets/hulx2002/VisDocAgentBench.CrossPoint-Bench
CrossPoint-Bench
CrossPoint-Bench is a comprehensive benchmark for evaluating Vision-Language Models (VLMs) on cross-view point correspondence tasks. It assesses models' abilities to spatial understanding, and correspondence between different viewpoints.
Dataset Structure
CrossPoint-Bench/
├── CrossPoint-Bench.jsonl # Main benchmark data file
└── image/
├── origin_image/ # Original scene images organized by scene ID
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/WangYipu2002/CrossPoint-Bench.FloodNet_2021-Track_2_Dataset_HF
FloodNet: High Resolution Aerial Imagery Dataset for Post-Flood Scene Understanding
This is the HF-hosted version of FloodNet.
The FloodNet 2021: A High Resolution Aerial Imagery Dataset for Post-Flood Scene Understanding provides high-resolution UAS imageries with detailed semantic annotation regarding the damages. To advance the damage assessment process for post-disaster scenarios, the authors of the dataset presented a unique challenge considering classification, semantic… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/FloodNet_2021-Track_2_Dataset_HF.embodied-spatial-reasoning
Embodied Spatial Reasoning Tasks
Dataset Description
This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.simverse2026
SimVerse
⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review.
A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.razavi-benchRazavi-bench
An expert-curated benchmark for analog-design reasoning.
Razavi-bench packages the question-answer assessments from Behzad Razavi's
Analog Design Experiments With AI Part 1 and Part 2 into a clean
one-task-per-directory benchmark. The tasks probe whether a model can reason
about MOS devices, small-signal circuits, feedback, oscillators, comparators,
dividers, LNAs, TIAs, and LC oscillators.
Each task directory keeps only the benchmark prompt, figure, and curated… See the full description on the dataset page: https://huggingface.co/datasets/Arcadia-2026/razavi-bench.Traffic-VQAgaokao-related-math
problems-scraper
爬取 出卷网 高考专区-数学试卷,转换为 JSONL 格式。
安装
pnpm install
用法
# 首次试跑 50 套
pnpm scrape:test
# 全量爬取 (3027 套,约 5-6 小时)
pnpm scrape:full
# 增量爬取 (每天最新)
pnpm scrape:resume
输出
data/
├── jsonl/problems.jsonl # 每行一个 JSON 题目
└── images/<id>/ # 试卷配图(散点图/几何图等本地副本)
state/
├── state.json # 爬取状态(maxSeenDate / fromDate)
└── seen.bin # 已抓试卷 ID 集合(断点续抓)
Cron(每日增量)
# 每天 3:00~6:00… See the full description on the dataset page: https://huggingface.co/datasets/cheese233/gaokao-related-math.thor25
Thor25: THORChain Cross-Chain Data
This repository contains the Thor25 dataset released with LOCARD: An Agentic Framework for Blockchain Forensics, published at IEEE ICBC 2026, Brisbane, Australia, June 1-5, 2026.
Thor25 supports research on agentic blockchain forensics, cross-chain transaction tracing, and evidence-grounded tool use by LLM agents. It does not map cleanly to conventional NLP task categories such as question answering or text generation: solving each benchmark… See the full description on the dataset page: https://huggingface.co/datasets/xhyumiracle/thor25.InternVL-Chat-V1-2-SFT-Data
Data Card for InternVL-Chat-V1-2-SFT-Data
Overview
Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT.
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.
