datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hleGameQA-140K
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
🎊 News
[2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.hle-extractoercommons-v1-optimized
OERCommons v1 Optimized
Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized
OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.VideoThinkBench
[CVPR 2026] Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
🎊 News
[2026.02] 🔥🔥Our work has been accepted by CVPR 2026! 🎉🎉🎉
[2025.11] Our paper "Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm" has been released on arXiv! 📄 [Paper] On HuggingFace, it has achieved "#1 Paper of the Day"!
[2025.11] 🔥We release "minitest" of our VideoThinkBench, including 500… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/VideoThinkBench.rendered-bookcorpus-bigrams2trial-v0-20250313
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/team-wonders/trial-v0-20250313.mathglyph-pages
MathGlyph Pages 1k With Detector Boxes
Synthetic mixed handwritten math pages prepared for the DFS/Rukopys detector pretraining pipeline.
Links
Generator repo: reirei-00/mathglyph_pages
Format
train/images/: train page images.
validation/images/: validation page images.
annotations/instance_train.json: COCO-style detector annotations for train.
annotations/instance_val.json: COCO-style detector annotations for validation.
train/metadata.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/dfs-team/mathglyph-pages.CC-Bench
CC-Bench
Structure
CC-Bench/
├── train/
│ ├── medical_diagnosis/
│ │ ├── imgs/
│ │ └── train.json
│ ├── industrial_inspection/
│ │ ├── imgs/
│ │ └── train.json
│ └── video_surveillance/
│ ├── imgs/
│ └── train.json
└── test/
├── medical_diagnosis/
│ ├── imgs/
│ ├── test_A.json
│ ├── test_B.json
│ ├── gt_A.json
│ └── gt_B.json
├── industrial_inspection/
│ ├── imgs/
│ ├── test_A.json… See the full description on the dataset page: https://huggingface.co/datasets/Megvii-Algo-Team/CC-Bench.lexgraph
Lexgraph dataset
The data plane of Lexgraph —
German legislation modelled as Laws as Git: a temporal, multi-authority
event log with HEAD, commits, open/closed branches and evidence-bound merges
(Bund / Bayern / EU; Länder records only after verification at the originating
Landtag). Built 2026-07-19.
Git is a navigation metaphor, not a substitute for legal status. Every row's
official source and status controls whether it is current law, a pending branch
or a documented… See the full description on the dataset page: https://huggingface.co/datasets/SNTIQ-Team/lexgraph.dataSpatial-Visualization-Benchmark
Spatial Visualization Benchmark
This repository contains the Spatial Visualization Benchmark. The evaluation code is released on: wangst0181/Spatial-Visualization-Benchmark.
Dataset Description
The SpatialViz-Bench aims to evaluate the spatial visualization capabilities of multimodal large language models, which is a key component of spatial abilities. Targeting 4 sub-abilities of Spatial Visualization, including mental rotation, mental folding, visual penetration, and… See the full description on the dataset page: https://huggingface.co/datasets/PLM-Team/Spatial-Visualization-Benchmark.rendered-wiki_en-bigramsinternal-v0-migrated-20260402-141930xlerobot-left-pick-cup-sep2-20260902-18193green-probe-dataset
Green Probe Dataset
Dataset collected to test YOLO stuff.
Remember to change the paths in config.yaml
data-imageGameQA-5KIn this repository, we specifically provide the 5k training samples from the complete GameQA-140K dataset used in our work for GRPO training of the models.
Refer to our paper for details. And our code for training and evaluation is at https://github.com/tongjingqi/Code2Logic.
Code2Logic: Game-Code-Driven Data Synthesis for Enhancing VLMs General Reasoning
This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-5K.xlerobot-left-pick-cup-sep3-20260903-2015SciCap-MLBCAP
MLBCAP: Multi-LLM Collaborative Caption Generation in Scientific Documents
📄 PaperMLBCAP has been accepted for presentation at AI4Research @ AAAI 2025. 🎉
📌 Introduction
Scientific figure captioning is a challenging task that demands contextually accurate descriptions of visual content. Existing approaches often oversimplify the task by treating it as either an image-to-text conversion or text summarization problem, leading to suboptimal results. Furthermore, commonly… See the full description on the dataset page: https://huggingface.co/datasets/TEAMREBOOTT-AI/SciCap-MLBCAP.coco-500
Dataset Card for "coco-500"
More Information needed
trial-v0-202504123D-STARESteamcraft_data
Dataset Card for TeamCraft
The TeamCraft dataset is designed to develop multi-modal, multi-agent collaboration in Minecraft. It features 55,000 task variants defined by multi-modal prompts and procedurally generated expert demonstrations.
This repository contains the data for the validation set and its visualizations.
To use the validation set, download TeamCraft-Data-Valid.zip and extract using unzip TeamCraft-Data-Valid.zip.
In addition, the training set is available in two… See the full description on the dataset page: https://huggingface.co/datasets/teamcraft/teamcraft_data.xlerobot-left-pick-cup-sep3-20260903-2054motogp_2025_teams_datasetdrone-panoramicForgeryVCR
ForgeryVCR Data
This repository stores the preprocessed 832-size assets used by the public
ForgeryVCR training and evaluation workflow.
datasets/
├── train/
│ ├── sft/ # CASIA v2 archives for expert selection and Agent SFT
│ └── rl/ # IMD2020 and FantasticReality archives for GRPO
└── test/ # public evaluation benchmark archives
Tampered-set JSON files use paths relative to their own dataset directory, for
example ./forged/example.png and ./gt/example_gt.png.… See the full description on the dataset page: https://huggingface.co/datasets/ForgeryVCR-Team/ForgeryVCR.
