datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mini-imagenet
Dataset Description
A mini version of ImageNet-1k with 100 of 1000 classes present.
Unlike some 'mini' variants this one includes the original images at their original sizes. Many such subsets downsample to 84x84 or other smaller resolutions.
Data Splits
Train
50000 samples from ImageNet-1k train split
Validation
10000 samples from ImageNet-1k train split
Test
5000 samples from ImageNet-1k validation split (all 50 samples per class)… See the full description on the dataset page: https://huggingface.co/datasets/timm/mini-imagenet.mind2web_multimodal_test_domain
Dataset Card for "Cross-Domain" Test Split in Multimodal Mind2Web
Note: This dataset is the test split of the Cross-Domain dataset introduced in the paper.
This is a FiftyOne dataset with 4050 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_domain.mind2web_multimodal_test_task
Dataset Card for Multimodal Mind2Web "Cross-Task" Test Split
Note: This dataset is the test split of the Cross-Task dataset introduced in the paper.
This is a FiftyOne dataset with 1338 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_task.Minecraft-Skins-20M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 19,973,928 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
id: A randomly generated UUID for each skin entry. These UUIDs are not linked to any external APIs or services (such as Mojang's player UUIDs) and serve solely as… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/Minecraft-Skins-20M.mini_imagenet
Dataset Card for "mini_imagenet"
More Information needed
mind2web_multimodal_test_website
Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split
Note: This dataset is the test split of the Cross-Website dataset introduced in the paper.
This is a FiftyOne dataset with 1019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.mini-reachy-animation
Reachy Mini Animation Dataset
Multi-view renders of 85 emotional animations performed by the Reachy Mini
robot, paired with the full robot joint state for every single frame.
Source of the animations. The emotional animations rendered here come from the
official pollen-robotics/reachy-mini-emotions-library
dataset by Pollen Robotics. This dataset re-renders those emotions from 12 camera
angles (with 3 background variants) and pairs every frame with the robot's joint state.… See the full description on the dataset page: https://huggingface.co/datasets/BastienATOS/mini-reachy-animation.Minecraft-Skins-Captioned-1M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical.
image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.gelsight-mini-pretrain
GelSight Mini Pretrain
~853K GelSight Mini tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92.
Frames
Sources
Real
536K
FoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad
Sim
317K
sim_tactile_mnist, sim_starstruck (Taxim-rendered, Mini-calibrated)
NC… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain.xyran_train_dataset
Xyran training dataset
Prepared for Xyran local content-safety model training.
Layout
sfw/
nsfw/
nsfl/
_manifests/
anime_dbrating WebP migration: 20260905_111217
Mapping:
general + sensitive -> SFW
questionable + explicit -> NSFW
WebP normalization:
existing WebP: byte-for-byte passthrough
non-WebP: libvips -> WebP
quality: 95
lossless: False
effort: 2
no resize
no crop
ZIP payload: STORE
Results:
SFW images: 683,275
NSFW images: 598,226
SFW… See the full description on the dataset page: https://huggingface.co/datasets/MingSafeR/xyran_train_dataset.arbiter-mini
Arbiter-mini
A small, purpose-built image dataset of household items captured under controlled Raspberry Pi camera conditions and labeled for binary waste/recycle classification according to San Diego, CA municipal recycling rules. Built as deployment-condition training data for the Arbiter sorting system, intended to be used alongside TrashNet to close the domain gap between studio imagery and real Pi-camera inference.
Motivation
Models trained purely on TrashNet… See the full description on the dataset page: https://huggingface.co/datasets/aaryavlal/arbiter-mini.UrbanPersona-120K-Interpretive
UrbanPersona-120K-Interpretive
Annotation corpora and analysis outputs for "Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation" (EMNLP 2026 Workshop Pandora). Two
multimodal LLMs, Qwen3-VL-8B and Gemma4 E4B, annotate the same 50 PerceptSent urban scenes as the
same 1,200 demographic personas at T = 0.1, 60,000 persona × image attempts per model and 120,000
in all, next to their no-persona ablations, a greedy T = 0 decoding… See the full description on the dataset page: https://huggingface.co/datasets/MInDS-lab-UTFPR/UrbanPersona-120K-Interpretive.CraftSight-Minecraft
CraftSight
Frame-level multi-label visual annotations for Minecraft agents.
Dataset creation toolkit: CraftSight is built and maintained with CraftSight Labeler, an open-source, browser-based annotation tool and reproducible release pipeline for multi-label Minecraft vision datasets. It supports manual and model-assisted labeling, structured game-state annotations, and trajectory-safe train/validation/test splits.
CraftSight provides Minecraft gameplay frames annotated… See the full description on the dataset page: https://huggingface.co/datasets/Krows7/CraftSight-Minecraft.minecraft-biomes
Minecraft Biomes (RGBD, pseudo-labeled)
Pseudo-labeled RGBD screenshots from Minecraft, covering 12 broad biome
categories. Each sample is an (RGB, depth) pair at 640×360 resolution.
Source
RGBD frames: 1908 paired (rgb, depth) samples from
zid8/syntheticMinecraftRGBD,
collected via MineRL.
Labels: Generated by a Gemma-3-4B model fine-tuned with LoRA on
the willowc/minecraft-biomes
dataset, then augmented with ~160 hand-selected MineRL ocean frames
to fix an… See the full description on the dataset page: https://huggingface.co/datasets/Wafik20/minecraft-biomes.minc-2500_split_1
Materials in Context Dataset (MINC-2500)
Dataset Summary
(from the website)
MINC-2500 is a patch classification dataset with 2500 samples per category
(Section 5.4 of the paper). This is a subset of MINC where samples have been
sized to 362 x 362 and each category is sampled evenly. The original resolution
images are not needed as we include the extracted patches in the archive.
gelsight-mini-gelsight-20260617
Dual GelSight Tactile Dataset 20260617
This dataset contains synchronized marker-mask tactile captures from two GelSight-style sensors:
gelsight_mini: GelSight Mini camera stream
gelsight: custom UVC GelSight-style camera stream
The data was collected on 2026-06-17 for real-world finetuning/adaptation of UniForce-style tactile models.
Directory Layout
marker/<sensor>/<indenter>/<frame_id>.jpg
collection_log.txt
The marker folders contain marker mask images… See the full description on the dataset page: https://huggingface.co/datasets/LancetRobotics/gelsight-mini-gelsight-20260617.UrbanPersona-60K
UrbanPersona-60K: persona-conditioned urban sentiment annotations
Every annotation produced for "Stable Behavior, Limited Variation: Persona Validity in LLM
Agents for Urban Sentiment Perception" (arXiv:2604.28048): 60,000 attempts in
which 1,200 demographically distinct LLM personas each judged the same 50 urban scenes, plus the
two no-persona ablations the paper measures them against, the seed profiles that produced the
personas, and the full analysis outputs.
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/MInDS-lab-UTFPR/UrbanPersona-60K.minang_foodMiniCAMELYON17
MiniCAMELYON — Patch-Level Breast Cancer Metastasis Detection
This dataset is derived from the CAMELYON17 challenge (CAncer MEtastases in LYmph nOdes challeNge), organized by the Diagnostic Image Analysis Group (DIAG) and the Department of Pathology of Radboud University Medical Center in Nijmegen, the Netherlands, as the second grand challenge in computational pathology. The original CAMELYON17 cohort comprises 1,000 H&E-stained whole-slide images of sentinel lymph node… See the full description on the dataset page: https://huggingface.co/datasets/chehablab/MiniCAMELYON17.mini-imagenet
Dataset Description
A mini version of ImageNet-1k with 100 of 1000 classes present.
Unlike some 'mini' variants this one includes the original images at their original sizes. Many such subsets downsample to 84x84 or other smaller resolutions.
Data Splits
Train
50000 samples from ImageNet-1k train split
Validation
10000 samples from ImageNet-1k train split
Test
5000 samples from ImageNet-1k validation split (all 50 samples per class)… See the full description on the dataset page: https://huggingface.co/datasets/Chronoglitter/mini-imagenet.mini-croupier
Dataset Description
TODO
Dataset Summary
TODO
Dataset Creatioon
TODO
open-images-v7-mini
AnnotateIt · Open the app · Models & datasets · Documentation
AnnotateIt Open Images V7 Mini Collection
Eight small, real-world, AnnotateIt-compatible datasets curated from the Open Images V7 validation split. Each archive contains 50–200 images, production-exported annotations, source details, and per-image attribution.
Only images whose official Open Images metadata lists CC BY 2.0 are included. Open Images annotations are CC BY 4.0. Open Images recommends independently… See the full description on the dataset page: https://huggingface.co/datasets/AnnotateIt/open-images-v7-mini.miniddbs-jpegThis repository contains the JPEG version of Mini-DDBS breast cancer dataset from Kaggle. Sinced the whole dataset is 50GB only the JPEG version is extracted and uploaded here for faster retrieving.
You can use Git Clone to download the whole data or the zip version (although zip version are recommended):
To access the data:
Using Git Clone
git clone https://huggingface.co/datasets/keanteng/miniddbs-jpeg
Using Python
from huggingface_hub import hf_hub_download
# Replace with the actual… See the full description on the dataset page: https://huggingface.co/datasets/keanteng/miniddbs-jpeg.minecraft-skins-1.1m-deduped-64x64
Minecraft Skins 1.1M Deduped (64x64 Edition)!
Nyuuzyou's Minecraft-Skins-20M but deduped using BLAKE3 hashes (only catches if pixel values are exactly the same), then filtered to only valid 64x64 skins.
Format is a 2.6 GB ZIP archive containing 64x64 PNG skin files.
Tools used
PIL Image (Python), Google Colab (free CPU tier) and BLAKE3 (Python)
How it was made
Loaded Nyuuzyou's Minecraft-Skins-20M,
Deduped using BLAKE3, only catching if pixel… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/minecraft-skins-1.1m-deduped-64x64.rvl_cdip_mini
Dataset Card for RVL-CDIP-MINI
This dataset is a subset (1%) of the original aharley/rvl_cdip merged with the corresponding annotations from jordyvl/rvl_cdip_easyocr.
You can easily and quickly load it:
dataset = load_dataset("dvgodoy/rvl_cdip_mini")
DatasetDict({
train: Dataset({
features: ['image', 'width', 'height', 'category', 'ocr_words', 'word_boxes', 'ocr_paragraphs', 'paragraph_boxes', 'label'],
num_rows: 3200
})
validation: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/dvgodoy/rvl_cdip_mini.mini-imagenet
Dataset Description
A mini version of ImageNet-1k with 100 of 1000 classes present.
Unlike some 'mini' variants this one includes the original images at their original sizes. Many such subsets downsample to 84x84 or other smaller resolutions.
Data Splits
Train
50000 samples from ImageNet-1k train split
Validation
10000 samples from ImageNet-1k train split
Test
5000 samples from ImageNet-1k validation split (all 50 samples per class)… See the full description on the dataset page: https://huggingface.co/datasets/Reyos/mini-imagenet.sentinel-lfm-mining-patches
sentinel-lfm — illegal-mining single-frame patches
128px RGB patches cropped from the Roboflow illegal-mining dataset, labelled
mine (1) / no-mine (0). Split by source image (no leakage) into
train/val/test. Provided as PNGs + vlm_sft-format JSONL (one image + prompt
-> JSON answer) so it drops straight into VLM fine-tuning.
split
pos
neg
total
train
1410
555
1965
val
303
66
369
test
303
116
419
RGB only (no multispectral). Each JSONL row is a single-turn VLM… See the full description on the dataset page: https://huggingface.co/datasets/ASTRALK/sentinel-lfm-mining-patches.SUN-minifashion-mnist-mini
AnnotateIt · Open the app · Models & datasets · Documentation
AnnotateIt Fashion-MNIST Mini
Small, deterministic, AnnotateIt-compatible samples derived from Fashion-MNIST.
Upstream revision: b2617bb6d3ffa2e429640350f613e3291e10b141Upstream license: MIT
File
Sample
Tasks
Format
Images
Size
SHA-256
Derived task
fashion-mnist-object-detection-mini.zip
Fashion-MNIST Detection Mini
Detection
COCO
80
0.08 MiB… See the full description on the dataset page: https://huggingface.co/datasets/AnnotateIt/fashion-mnist-mini.TFQ-Data-Full
TFQ-Data: A Fine-Grained Dataset for Image Implication
TFQ-Data is a large-scale visual instruction tuning dataset specifically designed to train Multi-modal Large Language Models (MLLMs) on Image Implication and Metaphorical Reasoning.
Unlike standard VQA datasets that focus on literal description, TFQ-Data utilizes a True-False Question (TFQ) format. This format provides high knowledge density and verifiable reward signals, making it an ideal substrate for Visual Reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/MING-ZCH/TFQ-Data-Full.
