datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MARS-Hyperspectral-EnMAP-PRISMA-v2025
MARS-Hyperspectral dataset (EnMAP and PRISMA) - version v2025
Updated version of the dataset. More information will be added soon.
FACETNayanaOCR_Corpus_2025
🪷 NayanaOCR Corpus 2025
A 1M-page, 22-language fully-parallel synthetic OCR + VQA corpus for document-centric vision-language models — every page rendered in every language.
NayanaOCR Corpus 2025 is one of the largest open-source multilingual, multi-task document datasets for training and evaluating OCR, layout detection, and visual question answering (VQA) in low-resource and underrepresented languages.
The headline property: it's a true parallel corpus. The same ~45,700 source… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/NayanaOCR_Corpus_2025.UNS-SSDS-2025Dataset from UNS SSDS 2025 Competition
AIC20252025-10-16T21-42-02plus00-00_gdpvalThe original repo ID for this repo is jeqcho/2025-10-16T21-42-02plus00-00_gdpval.
The repo is the result from running GPT-5 with thinking low on the public subset of GDPval implemented in Inspect Eval, which has 220 tasks.
The goal was to test and validate the Inspect Eval implementation of GDPval.
The repo was uploaded in 16 October 2025, 21:42:02 UT+0 and submitted to OpenAI's GDPval grading server shortly after.
OpenAI got back to us on December 2 with the results. Note that their system… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/2025-10-16T21-42-02plus00-00_gdpval.kangaroo_2025_5_6
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 5-6 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_5_6.kangaroo_2025_7_8
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 7-8 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_7_8.kangaroo_2025_9_10
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 9-10 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_9_10.kangaroo_2025_3_4
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 3-4 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_3_4.kangaroo_2025_1_2
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 1-2 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_1_2.kangaroo_2025_11_12
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 11-12 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_11_12.danbooru2025-metadata
🎨 Danbooru 2025 Metadata
Latest Post ID: 9,158,800
(as of Apr 16, 2025)
📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork.
Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes:
More consistent tag history tracking
Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.MATPMATPBench
visual_ai_at_neurips2025
Dataset Card for neurips-2025-vision-papers
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/visual_ai_at_neurips2025")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/visual_ai_at_neurips2025.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.SN-GSR-2025
SoccerNet Challenge 2025 - Game State Reconstruction
Download the dataset
Install the huggingface_hub pip package:
pip install huggingface_hub[cli]
Download the dataset with the following Python code :
from huggingface_hub import snapshot_download
snapshot_download(repo_id="SoccerNet/SN-GSR-2025",
repo_type="dataset", revision="main",
local_dir="SoccerNet/SN-GSR-2025")
SHROOMCAP2025_DataSet
🍄 SHROOM-CAP 2025 Unified Hallucination Detection Dataset
📖 Dataset Description
This dataset was created for the SHROOM-CAP 2025 Shared Task on multilingual scientific hallucination detection. It combines and unifies multiple hallucination detection datasets into a single, balanced training corpus for fine-tuning XLM-RoBERTa-Large.
Key Features
124,821 total samples - 172x larger than original SHROOM training data
Perfectly balanced - 50% correct… See the full description on the dataset page: https://huggingface.co/datasets/Haxxsh/SHROOMCAP2025_DataSet.CC3Msakugabooru2025
Sakugabooru2025: Curated Animation Clips from Enthusiasts
Sakugabooru.com is a booru-style imageboard dedicated to collecting and sharing noteworthy animation clips, emphasizing Japanese anime but open to creators worldwide. Over the years, it has amassed more than 240,000 animation clips, alongside informative blog posts for anime fans everywhere.
With the growing interest in generative video models and AI animations, the scarcity of proper animation-related video datasets has… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/sakugabooru2025.20250704_beifen_2conditionskangaroo_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
competition (string): Competition or… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025.VanGogh_vs_TreeOilPainting_QuantumTorque_EnergyField_Analysis_Phase1_2025
Dataset Policy
VanGogh Vs. Tree Oil Painting: Quantum Torque Energy Field Analysis 2025
Structure Type
Free-form and Semi-structured Narrative
Core Principles
Each file is an independent analytical entity with its own identity.
Each file is the result of Autonomous AI–Human Co-analysis.
The structure is intentionally open, flexible, and adaptive, reflecting the natural reasoning process of the researcher, rather than forcing rigid… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/VanGogh_vs_TreeOilPainting_QuantumTorque_EnergyField_Analysis_Phase1_2025.cmp-vialitix-photos-2025
Panoramax sig14 photos
Attribution — Ce jeu de données contient des images issues de la plateforme
Panoramax de l'IGN (Institut national de l'information
géographique et forestière). Les images d'origine sont publiées sous
Licence Ouverte / Open License 2.0 (Etalab)
par leur producteur sig14 (Service d'Information Géographique du Calvados) et l'IGN.
Street-view photos from the IGN Panoramax instance, scraped from the user sig14
(Service d'Information Géographique du Calvados… See the full description on the dataset page: https://huggingface.co/datasets/calvadosdep/cmp-vialitix-photos-2025.trash-in-river-2025
Street Parade 2025 Dataset
Overview
This dataset was collected by SARA, a student initiative at ETH Zurich, to enable open research on trash presence in aquatic environments. It contains images of litter in the Limmat River in Zurich the day after the Street Parade (August 9, 2025). The dataset is intended for training and evaluating trash classification models.
Dataset summary
Collection date: August 9, 2025
Location: Kornhausbrücke, Zurich… See the full description on the dataset page: https://huggingface.co/datasets/SARA-smartphone-assisted-river-analysis/trash-in-river-2025.ai-intern-challenge-2025Anonymous_ACMMM_2025_Submission
🗂️ Anonymous_ACMMM_2025_Submission Dataset
This dataset is prepared for the Anonymous ACMMM 2025 submission, containing multi-view event-based data designed for dynamic 3D scene reconstruction tasks.
📁 Dataset Structure
Each subfolder corresponds to a distinct synthetic or real-world scene, such as:
lego_6_views/
capsule_6_views/
garage_6_views/
Restroom_6_views/
Cubes_6_views/
Hinge_6_views/
MC-Toy_6_views/
Rubik’s-Cube_6_views/
Each scene folder contains 6 views… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-ACMMM-2025-Submission/Anonymous_ACMMM_2025_Submission.enginuity-bench
Enginuity Bench
Enginuity Bench is a multimodal benchmark dataset built from U.S. military vehicle repair parts manuals (Technical Manuals, TM series). It is designed to evaluate vision-language models on two tasks grounded in real-world engineering documentation:
Component Identification — given a figure and its associated parts list table, identify and extract structured parts data
Question Answering — given a figure and its associated parts list, answer natural-language… See the full description on the dataset page: https://huggingface.co/datasets/enginuity2025/enginuity-bench.spatialva-runtime-assetssmart-product-pricing-2025
