datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hot3d
HOT3D-Clips
This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset.
Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here.
See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.).
More details can be found in the HOT3D paper and BOP 2024 report.
wds_objectnetMultiBanana-Benchmark🍌 MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 🍌
CVPR 2026 (Main)
This repository provides the datasets for
“MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation” by Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo and Hiroki Furuta
Paper Link
https://arxiv.org/abs/2511.22989
Github Repository
For the usage of this benchmark, please see Github… See the full description on the dataset page: https://huggingface.co/datasets/kohsei/MultiBanana-Benchmark.wds_imagenet_sketchPDE_Inverse_Problem_Benchmarking
PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems
This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems.
Code: GitHub - ASK-Berkeley/PDEInvBench
Sample Usage
You can use the provided script from the codebase to batch download the data:
pip install huggingface_hub
python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.BLINK
BLINK: Multimodal Large Language Models Can See but Not Perceive
🌐 Homepage | 💻 Code | 📖 Paper | 📖 arXiv | 🔗 Eval AI
This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive"
Introduction
We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans “within a… See the full description on the dataset page: https://huggingface.co/datasets/BLINK-Benchmark/BLINK.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.benchmark
SLM Lab
Modular Deep Reinforcement Learning framework in PyTorch.
Companion library of the book Foundations of Deep Reinforcement Learning.
Documentation · Benchmark Results
NOTE: v5.0 updates to Gymnasium, uv tooling, and modern dependencies with ARM support - see CHANGELOG.md.
Book readers: git checkout v4.1.1 for Foundations of Deep Reinforcement Learning code.
BeamRider
Breakout
KungFuMaster
MsPacman
Pong
Qbert
Seaquest
Sp.Invaders… See the full description on the dataset page: https://huggingface.co/datasets/SLM-Lab/benchmark.wds_imagenet-rGAIA
GAIA dataset
GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc).
We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format.
Data and leaderboard
GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.wds_imagenet-aGOAI-2026wds_imagenetv2oxford_flowers102RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.wds_imagenet1kAwesome_Spatial_VQA_Benchmarksjee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
🏆 View the live leaderboard → — interactive results across JEE Advanced, JEE Main & NEET, with open/closed-weight badges, contamination flags, and per-run cost.
A benchmark for evaluating vision-capable LLMs on Indian competitive exam questions (JEE Advanced & NEET). Each question is the original exam image; models answer via the OpenRouter API and are scored with authentic, exam-specific marking schemes — including partial credit for JEE… See the full description on the dataset page: https://huggingface.co/datasets/Reja1/jee-neet-benchmark.MOVA_benchmark_for_arena
MOVA Benchmark for Arena
This is the benchmark used for the subjective arena experiments of MOVA (MOVA: Towards Scalable and Synchronized Video–Audio Generation). All prompts are rewritten by the workflow introduced in the paper.
Paper: MOVA: Towards Scalable and Synchronized Video–Audio Generation
Code: https://github.com/OpenMOVA/MOVA
Overview
The benchmark contains 732 samples in total, organized into two subsets:
Subset
Samples
MOVA-Bench
132… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuzhang-0212/MOVA_benchmark_for_arena.IDEAL-Scenes
IDEAL-Bench: Indoor Dataset for Evaluating Analysis by 3D Layout Reasoning
IDEAL-Bench is an evaluation suite that requires VLMs to predict structured 3D layouts on photorealistic indoor scenes across 10 room types, scored along five numerical dimensions (scene validity, physical plausibility, geometric accuracy, object recognition, and grid layout) and a perceptual render-and-compare protocol.
Built on IDEAL-Scenes - 1,000 procedurally generated, re-renderable Blender scenes… See the full description on the dataset page: https://huggingface.co/datasets/IDEAL-Benchmark/IDEAL-Scenes.ocr-benchmark
OmniAI OCR Benchmark
A comprehensive benchmark that compares OCR and data extraction capabilities of different multimodal LLMs such as gpt-4o and gemini-2.0, evaluating both text and JSON extraction accuracy.
Benchmark Results (Feb 2025) | Source Code
wds_fer2013benchmark-datasets
Latency-Sensitive Bench datasets
Accepted zero-latency teacher rollouts for the supported benchmark tasks.
Viewer subsets
humanoidbench_balance_simple: 90 training and 10 validation episodes. The
observation.image values are PNG bytes declared as the Hugging Face Image
feature, so the Dataset Viewer renders them instead of showing their encoded
representation. Canonical LeRobot MP4 files remain under each split's
videos/ directory.
mikasa_intercept_grab_fast:… See the full description on the dataset page: https://huggingface.co/datasets/latency-sensitive-bench/benchmark-datasets.figma-slide-benchmark
Figma Slide Editing Benchmark
Benchmark accompanying our EMNLP 2026 Industry Track (Main) accepted paper "ACE: A
Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation
Automation".
📄 Paper: https://arxiv.org/pdf/2608.24103
💻 Code: https://github.com/BloomBerry/agentic-canvas-editor
Overview
Each benchmark item is a slide-editing task defined as a pair of Figma Slides
documents:
*_TestA — the input deck the agent starts from.
*_GroundTruthA — the… See the full description on the dataset page: https://huggingface.co/datasets/BloomBerry/figma-slide-benchmark.UAVDT-Benchmark-Mgraphmemix-benchmarks
GraphMemix Benchmarks
Unified multimodal memory benchmark bundles used by
GraphMemix
(arXiv:2608.26983) — four long-term
personalized memory benchmarks with their raw media assets, packaged together
for reproducible evaluation.
Benchmark
Questions
Memories
Track
Upstream license
ATM-Bench (default + hard)
1,044
11,034
memory QA over one multimodal archive
MIT
Mem-Gallery
1,711
7,944
multimodal gallery memory QA
MIT
MemEye
1,855
3,392
comics-derived memory QA… See the full description on the dataset page: https://huggingface.co/datasets/oking0197/graphmemix-benchmarks.Food_Portion_Benchmark
Food Portion Benchmark (FPB) Dataset
The Food Portion Benchmark (FPB) is a comprehensive dataset and benchmark suite for multi-task food scene understanding, combining food detection and portion size (weight) estimation. It was introduced to support research in dietary analysis, nutrition tracking, and food computing. The dataset is built with high-quality annotations and evaluated using an extended YOLOv12-based multi-task model .
📦 Dataset Overview
Total images:… See the full description on the dataset page: https://huggingface.co/datasets/issai/Food_Portion_Benchmark.benchmarkwds_flickr30kCharge-050_0130
