datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coco-2017-mirror
COCO 2017 mirror
This is a just mirror of the raw COCO dataset files, for convenience. You have to download it using something like:
pip install huggingface_hub
huggingface-cli download --local-dir coco-2017 pcuenq/coco-2017-mirror
And then unzip the files before use.
nerf-synthetic-mirrorLPNSR
LPNSR Dataset
This repository contains the evaluation datasets and testing data associated with the paper LPNSR: Optimal Noise-Guided Diffusion Image Super-Resolution Via Learnable Noise Prediction.
Project Links
Paper: arXiv:2603.21045
GitHub Repository: Faze-Hsw/LPNSR
Dataset Description
This dataset collection is used to evaluate image super-resolution models on both synthetic and complex real-world degradations. It contains pairs of Low-Quality (LQ) and… See the full description on the dataset page: https://huggingface.co/datasets/mirpri/LPNSR.MirrorPPR47MPaper: https://arxiv.org/abs/2606.29308
mirage18k
[IROS 2026] Mirage 18k: Dataset for Glass Segmentation & Depth Estimation
Mirage 18k is a novel, multi-task dataset comprising 18,353 manually annotated images across 38 unique indoor scenes, designed specifically for joint glass segmentation and glass-aware monocular depth estimation in robotics.
It contains diverse real-world glass structures (indoor panes, frosted doors, windows, clear doors) with severe background clutter, saliency, and dynamic obstacles.
Model Checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/rtarun1/mirage18k.mb-frost_cls
mb-frost_cls
A Mars image classification dataset for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-14
Cite As: TBD
Classes
This dataset contains the following classes:
0: frost
1: non_frost
Statistics
train: 30124 images
test: 12249 images
val: 11415 images
few_shot_train_2_shot: 4 images
few_shot_train_1_shot: 2 images
few_shot_train_10_shot:… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-frost_cls.anime-syntheticsMostly unfiltered anime-style images generated by various text to image models, collected from various sources (some were submitted for inclusion by their creators).
Includes a subset of p1atdev/niji-v5, albeit captioned differently than the source.
Contains 2224 image & caption pairs.
As it is unfiltered, some adult content may be included.
Captions may not be completely accurate.
If you wish to submit content, do it as a pull request.
miracl-vision
MIRACL-VISION
MIRACL-VISION is a multilingual visual retrieval dataset for 18 different languages. It is an extension of MIRACL, a popular text-only multilingual retrieval dataset. The dataset contains user questions, images of Wikipedia articles and annotations, which article can answer a user question. There are 7898 questions and 338734 images. More details can be found in the paper MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark.
This dataset is ready… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/miracl-vision.Mirage-Test
🌊 Mirage-Test Dataset
Mirage-Test is a modern test-only dataset for benchmarking AI-generated image detection models.
It contains real (0_real) and fake (1_fake) images across five distinct content domains, designed to evaluate generalization across diverse visual semantics.
The fake images are generated using state-of-the-art generative models specifically optimized for perceptual realism and visual fidelity.
📌 This dataset is for evaluation only. No training split is… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/Mirage-Test.MIRAGE
MIRAGE Benchmark
Project Page | Paper | GitHub
MIRAGE is a benchmark for multimodal expert-level reasoning and decision-making in consultative interaction settings, specifically designed for the agriculture domain. It captures the complexity of expert consultations by combining natural user queries, expert-authored responses, and image-based context.
The benchmark spans diverse crop health, pest diagnosis, and crop management scenarios, including more than 7,000 unique biological… See the full description on the dataset page: https://huggingface.co/datasets/MIRAGE-Benchmark/MIRAGE.MIRA
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
Dataset Description
MIRA (Multimodal Imagination for Reasoning Assessment) evaluates whether MLLMs can think while drawing—i.e., generate and use intermediate visual representations (sketches, diagrams, trajectories) as part of reasoning.MIRA includes 546 carefully curated problems spanning 20 task types across four domains:
Euclidean Geometry (EG)
Physics-Based Reasoning (PBR)… See the full description on the dataset page: https://huggingface.co/datasets/YiyangAiLab/MIRA.mb-domars16k
mb-domars16k
A Mars image classification dataset for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-14
Cite As: TBD
Classes
This dataset contains the following classes:
0: ael
1: rou
2: cli
3: aec
4: tex
5: smo
6: fss
7: rid
8: fse
9: sfe
10: fsf
11: fsg
12: sfx
13: cra
14: mix
Statistics
train: 11305 images
test: 1614 images
val: 3231 images… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-domars16k.DigiCam-Mirflickr-MultiMask-10Keuroc-mirrormirage_mvtec_visamb-crater_multi_seg
mb-crater_multi_seg
A segmentation dataset for planetary science applications.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-15
Cite As: TBD
Classes
This dataset contains the following classes:
0: Background
1: Other
2: Layered
3: Buried
4: Secondary
Directory Structure
The dataset follows this structure:
dataset/
├── train/
│ ├── images/ # Image files
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-crater_multi_seg.mb-landmark_cls
mb-landmark_cls
A Mars image classification dataset for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-14
Cite As: TBD
Classes
This dataset contains the following classes:
0: oth
1: cra
2: ddu
3: sst
4: bdu
5: ime
6: sch
7: spi
Statistics
train: 6997 images
test: 1793 images
val: 2025 images
few_shot_train_2_shot: 16 images… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-landmark_cls.mb-atmospheric_dust_cls_rdr
mb-atmospheric_dust_cls_rdr_upd
A Mars image classification dataset for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-22
Cite As: TBD
Classes
This dataset contains the following classes:
0: dusty
1: not_dusty
Statistics
train: 9817 images
test: 5214 images
val: 4969 images
few_shot_train_2_shot: 4 images
few_shot_train_1_shot: 2 images… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-atmospheric_dust_cls_rdr.Pocket-Rocket-1.0-Mid-Training-23MThis repository stores converted raw image WebDataset tar shards from multiple source datasets for streaming training.
mirainikki
Bangumi Image Base of Mirai Nikki
This is the image base of bangumi Mirai Nikki, we detected 27 characters, 2067 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mirainikki.mb-conequest_seg
mb-conequest_seg
A segmentation dataset for planetary science applications.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-15
Cite As: TBD
Classes
This dataset contains the following classes:
0: Background
1: Cone
Directory Structure
The dataset follows this structure:
dataset/
├── train/
│ ├── images/ # Image files
│ └── masks/ # Segmentation masks
├── val/… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-conequest_seg.mb-crater_binary_seg
mb-crater_binary_seg
A segmentation dataset for planetary science applications.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-15
Cite As: TBD
Classes
This dataset contains the following classes:
0: Background
1: Crater
Directory Structure
The dataset follows this structure:
dataset/
├── train/
│ ├── images/ # Image files
│ └── masks/ # Segmentation masks… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-crater_binary_seg.MIRB
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
File Structure
├── MIR
|── analogy.json
│── codeu.json
|── dataset_namex.json
└── Images
├── analogy
│ └── image_x.jpg
└──codeu
└── image_x.jpg
JSON Structure
{
"questions": " What is the expected kurtosis of the sequence created by`create_number_sequence(-10, 10)`?\n\n1.… See the full description on the dataset page: https://huggingface.co/datasets/VLLMs/MIRB.libero_spatial_onlyThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 432,
"total_frames": 52970,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:432"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mirage415/libero_spatial_only.MIRB-hfMira-Scene-Dataset
Mira-Scene Benchmark
The Mira-Scene Dataset repository contains the public evaluation benchmark
used to measure the Mira-Scene single-image 3D scene reconstruction pipeline.
The current release contains the BlendSwap evaluation scenes, with rendered
scene images, per-object instance masks, depth data, and ground-truth 3D
annotations. The Mira-Scene training data will be added to this same dataset
repository in a future release.
This dataset is intended for evaluation and… See the full description on the dataset page: https://huggingface.co/datasets/Yang-Tian/Mira-Scene-Dataset.DigiCam-Mirflickr-MultiMask-25KDigiCam-Mirflickr-MultiMask-1KMiroEval-data
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
MiroEval is a comprehensive evaluation framework for Deep Research systems, providing automated task generation and assessment across three complementary dimensions: Factual correctness, Point-wise quality, and Process quality.
Quick Start
1. Setup
All three evaluation modules share a single Python environment managed by uv at the repo root:
uv sync
If you use… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroEval-data.mirage-news
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
[Paper]
[Github]
This dataset contains a total of 15,000 pieces of real or AI-generated multimodal news (image-caption pairs) -- a training set of 10,000 pairs, a validation set of 2,500 pairs, and five test sets of 500 pairs each. Four of the test sets are out-of-domain data from unseen news publishers and image generators to evaluate detector's generalization ability.
=== Data Source (News Publisher + Image Generator)… See the full description on the dataset page: https://huggingface.co/datasets/anson-huang/mirage-news.
