datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.EMMOE-100
EMMOE-100 Trainset
Resources
Project
Paper
Code
Model
Dataset
Dataset Feature
Task Attributes
Task Example
Dataset Structure
EMMOE-100/
├── README.md
├── assets/
├── data/
│ └── train/
│ ├── 1/
│ │ ├── info.txt
│ │ ├── info_re1.txt
│ │ ├── info_re2.txt
│ │ ├── info_re3.txt
│ │ ├── keypath.json
│ │ ├── scene.json
│ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Dongping-Li/EMMOE-100.EarthDial-Dataset
🌍 EarthDial-Dataset
The EarthDial-Dataset is a curated collection of evaluation-only datasets focused on remote sensing and Earth observation downstream tasks. It is designed to benchmark vision-language models (VLMs) and multimodal reasoning systems on real-world scenarios involving satellite and aerial imagery.
📚 Key Features
Evaluation-focused: All datasets are for inference/testing only — no train/val splits.
Diverse Tasks:
Classification
Object Detection
Change… See the full description on the dataset page: https://huggingface.co/datasets/akshaydudhane/EarthDial-Dataset.EASI-Leaderboard-Data
EASI Leaderboard Data
A consolidated dataset for the EASI Leaderboard, containing the evaluation data (inputs/prompts) actually used on the leaderboard across spatial reasoning benchmarks for VLMs.
Looking for the Spatial Intelligence leaderboard?https://huggingface.co/spaces/lmms-lab-si/EASI-Leaderboard
🔎 Dataset Summary
Question types: MCQ (multiple choice) and NA (numeric answer).
File format: TSV only.
Usage: These TSVs are directly consumable by the EASI… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-si/EASI-Leaderboard-Data.SciMDR-EvalETCHR-SFT-400K
ETCHR SFT-400K
📖Paper
| 🏠Homepage
| 🤗ETCHR-FLUX.2-klein-9B Model
| 🤗ETCHR SFT-400K Dataset
| 🤗ETCHR GRPO-10K Dataset
| 🤗DL3DV-2K Benchmark
ETCHR SFT-400K is the SFT training data for transfering a passive instruction-following image editor (built on FLUX.2-klein-base-9B) into an autonomous, question-conditioned visual reasoning assistant. It contains 400,000 samples of five tasks (Fine-grained Perception, Chart Understanding, Maze Solving, Jigsaw Puzzle and… See the full description on the dataset page: https://huggingface.co/datasets/BeichenZhang/ETCHR-SFT-400K.taiwan-examsMachine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/taiwan-exams.EgoNormia
EgoNormia: Benchmarking Physical-Social Norm Understanding
MohammadHossein Rezaei*,
Yicheng Fu*,
Phil Cuvin*,
Caleb Ziems,
Yanzhe Zhang,
Hao Zhu,
Diyi Yang,
🌎Website |
🤗 Dataset |
📄 arXiv |
📄 HF Paper
EgoNormia
EgoNormia is a challenging QA benchmark that tests VLMs' ability to reason over norms in context.
The datset consists of 1,853 physically grounded egocentric
interaction clips from Ego4D… See the full description on the dataset page: https://huggingface.co/datasets/open-social-world/EgoNormia.Robot-EQ
RobotEQ-Data
Official dataset release for RobotEQ.
Evaluation & Scripts
For inference scripts, evaluation scripts, and data production tooling, see the RobotEQ code repository.
Dataset Statistics
Item
Count
Behavior judgment scenarios (synthetic)
1,812
Behavior judgment scenarios (real POV)
223
Behavior judgment scenarios (total)
2,035
Behavior judgment behavior annotations
3,171
Spatial grounding questions
825… See the full description on the dataset page: https://huggingface.co/datasets/Tongji-Emotion/Robot-EQ.LLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.Mem-Gallery
📖 Overview
Mem-Gallery is a comprehensive benchmark dataset designed to evaluate multimodal long-term memory capabilities of MLLM agents across multi-session conversations. The dataset features realistic, persona-driven dialogues spanning 20 scenarios, each enriched with contextual images to test memory retention, recall, and reasoning over extended interactions.
🎯 Key Features
Diverse Scenarios: Covering topics from AI & Robotics to Daily Life… See the full description on the dataset page: https://huggingface.co/datasets/Ethan-Bei/Mem-Gallery.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.Omni-Edu
Omni-Edu — Core V6 SFT mixture
69,999 supervised instruction examples (~158M characters) covering K-12 subject
competence, curriculum grounding, diagnostic reasoning, pedagogical action and
general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image
referenced by the JSONL ships in this repository under images/.
This is the system-prompted assembly of the v6 core mixture: every row carries
an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.ttrpg-rpg-fandom-com-en
ttrpg-rpg-fandom-com-en (RPG Fandom EN Dataset)
[Russian version below / Русская версия ниже]
Description
This dataset contains a complete dump of the English rpg.fandom.com wiki, converted to clean Markdown. It is designed for RAG (Retrieval-Augmented Generation), LLM fine-tuning, and research.
Structure
markdown/: Cleaned documents with metadata.
indexes/documents.jsonl: Global document registry.
indexes/chunks.jsonl: Semantic fragments for… See the full description on the dataset page: https://huggingface.co/datasets/exnihilum/ttrpg-rpg-fandom-com-en.muddle-eval-bundle-img
MUDDLE eval bundle - IMAGE modality
Each page rendered as one 150-DPI PNG. Page images are stored ONCE per distinct document
under images/<doc>/pNNNN.png and referenced by each cell question.json (image_docs,
in load order: source first, then distractors). 1350 cells, all hard negatives in [10,40] pages.
Related
Part of MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects
(COLM 2026 Workshop on Context Beyond the Window).
Piece… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/muddle-eval-bundle-img.EarthVLSetEarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
by Junjue Wang,
Yanfei Zhong,
Zihang Chen,
Zhuo Zheng,
Ailong Ma, and Liangpei Zhang
[Paper],
[Dataset]
News
2026/01/06, New Global-LoveDA !!! We released the global-scale segmentation leaderboard at Global-LoveDA. Just zip all the test images into one file and submit it.
2026/01/06, The segmentation data is released at [Dataset].
2026/01/06, We are preparing the code and data for… See the full description on the dataset page: https://huggingface.co/datasets/Kingdrone-Junjue/EarthVLSet.Mantis-Eval
Overview
This is a newly curated dataset to evaluate multimodal language models' capability to reason over multiple images. More details are shown in https://tiger-ai-lab.github.io/Mantis/.
Statistics
This evaluation dataset contains 217 human-annotated challenging multi-image reasoning problems.
Leaderboard
We list the current results as follows:
Models
Size
Mantis-Eval
LLaVA OneVision
72B
77.60
LLaVA OneVision
7B
64.20
GPT-4V
-
62.67… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Eval.MM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.EVADE-Bench
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
🤗 Dataset | Paper | GitHub
E-commerce platforms increasingly rely on Large Language Models and Vision-Language Models to detect illicit or misleading product content. However, these models remain vulnerable to evasive content, which refers to inputs that superficially comply with platform policies while covertly conveying prohibited claims. Unlike traditional adversarial attacks that aim to… See the full description on the dataset page: https://huggingface.co/datasets/koenshen/EVADE-Bench.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.jee-main-questions
JEE Main — Question Bank
A structured dataset of JEE Main examination questions with full metadata,
worked solutions, and diagrams. Built for education, ML training, and
question-generation use cases.
Subsets:
Chemistry — 738 questions from 28 papers
Physics — 768 questions from 28 papers
Mathematics — 801 questions from 28 papers
Over 2,300 questions across the three core JEE subjects.
Structure
Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-main-questions.EgoGazeVQA-91-nips25DB
EgoGazeVQA-91 • NeurIPS 2025 Datasets & Benchmarks submission
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
1 Folder layout
EgoGazeVQA-91-nips25DB/
├── qa\_pairs/ # VQA supervision
│ ├── causal\_ego4d.{csv,json}
│ ├── spatial\_ego4d.{csv,json}
│ ├── temporal\_ego4d.{csv,json}
│ └── ...
├── keyframe_tar/
│ ├── ego4d.tar.gz
│ ├── egoexo.tar.gz
│ └── egtea.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/anonupload/EgoGazeVQA-91-nips25DB.AVI-Math
Dataset Sources
Repository: https://github.com/VisionXLab/avi-math
Paper: https://arxiv.org/abs/2509.10059
BibTeX:
@ARTICLE{zhou2025avimath,
author={Zhou, Yue and Feng, Litong and Lan, Mengcheng and Yang, Xue and Li, Qingyun and Ke, Yiping and Jiang, Xue and Zhang, Wayne},
journal={ISPRS Journal of Photogrammetry and Remote Sensing},
title={Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/erenzhou/AVI-Math.aya-mm-exams-spanish-nursingNursing Spanish Exams for the Multimodal Aya Exams Projects.
Questions available in file: data.json
Images stored in: /images
Original data and file available here: link
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.data-estate
Quick start — get everything in one command
pip install -U huggingface_hub
hf download EMTIAZZ/data-estate --repo-type dataset --local-dir ./data-estate
This downloads everything (all tables, PDFs, scanned images, emails, and transcripts) into
./data-estate. See How to download for more ways.
Hints
These are hints, not answers. They point out what to look for in each kind of data and what
to think about. Picking the tools and building the pipeline is your job.… See the full description on the dataset page: https://huggingface.co/datasets/EMTIAZZ/data-estate.cc_news_pt_v2
Dataset Summary
This version of the dataset is the portuguese subset from stanford-oval/ccnews.
CC-News-PT v2 is a curation of +11 million news articles from CommonCrawl News in the Portuguese language, from the beginning (2016) to June of 2024.
The data has been cleaned and deduplicated, and language of articles have been detected and filtered. The process is similar to what HuggingFace's DataTrove does.
For license information, please refer to CommonCrawl's Terms of Use.… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/cc_news_pt_v2.MMS-e
MMS-e: Benchmarking the Resilience of Large Multimodal Models to Visual Scrambling
Benchmark Examples
Patchwise Question Answering: Divide the images into 2x2, 4x4, and 8x8 patches, then shuffle all the patches, and measure the ability of LMMs to answer questions about these images.
Reconstruction task: Let LMMs reconstruct the order of shuffled patches based on the image' s caption, and let LMMs reconstruct the shuffled caption based on the image.
Fixed Patch… See the full description on the dataset page: https://huggingface.co/datasets/jyjyjyjy/MMS-e.OpenClaw-EvalMix
OpenClaw EvalMix
OpenClaw EvalMix is a Harbor-format collection of 360 agent-evaluation tasks across four task families. Each task directory includes task.toml, an instruction, an environment definition, and verifier tests.
Composition
Family
Tasks
Local payload
clawbench
19
0.00 GiB
liveclawbench
134
0.07 GiB
pinchbench
147
0.02 GiB
wildclawbench
60
14.05 GiB
The repository preserves each family at the root so task paths remain direct.… See the full description on the dataset page: https://huggingface.co/datasets/zt1106/OpenClaw-EvalMix.
