datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-testtiny-imagenet
Dataset Card for tiny-imagenet
Dataset Summary
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images. Each class has 500 training images, 50 validation images, and 50 test images.
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=64x64 at 0x1A800E8E190,
'label': 15
}… See the full description on the dataset page: https://huggingface.co/datasets/zh-plus/tiny-imagenet.product-photography-v1-tiny-prompts-tasks-collage-filteredemnist-letters-tiny
Dataset Card for EMNIST-Letters-10k
A random subset of the train and test splits from the letters portion of EMNIST
This is a FiftyOne dataset with 10000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/emnist-letters-tiny.Tiny-GenImage
Tiny GenImage Dataset
📝 Dataset Description
Dataset Summary
The Tiny GenImage Dataset is a curated, scaled-down collection of images and associated metadata designed to train, validate, and benchmark models for detecting and identifying artificially generated content. The dataset contains a mix of real-world images alongside those generated by prominent AI models, including various diffusion models (like Stable Diffusion 1.4/1.5, GLIDE, Midjourney, ADM, VQDM… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/Tiny-GenImage.text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
TinyMixtral-4x248M-MoE-atlas
juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas
A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing.
If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.MindCube-TinyBench
MindCube TinyBench Evaluation Dataset
This package contains the MindCube TinyBench split used for MindCube evaluation in PhysBrainEvalKit. It is a compact evaluation-only subset for reproducible testing of vision-language models on spatial mental modeling tasks.
Contents
MindCube-TinyBench/
├── data/
│ ├── raw/
│ │ └── MindCube_tinybench.jsonl
│ └── other_all_image/
│ └── <referenced image files>
└── README.md
The split contains 1,050 questions and… See the full description on the dataset page: https://huggingface.co/datasets/VLyb/MindCube-TinyBench.tiny-audits
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DivyaApp/tiny-audits.conceptual_captions_3m_en_tinyneuralatlas-attributions-convnext_tiny
Neural Atlas attributions — convnext_tiny on imagenet-pico
Precomputed attribution maps and faithfulness metrics for the torchvision
convnext_tiny model (default pretrained weights, no fine-tuning) on imagenet-pico,
a 3000-image subset of ImageNet-1k with three images for each of the 1000
classes.
This repository is part of Neural Atlas, a web tool for comparing
attribution methods across vision architectures on the same image, developed
as an undergraduate thesis at the… See the full description on the dataset page: https://huggingface.co/datasets/Matgc04/neuralatlas-attributions-convnext_tiny.SPAR-Bench-Tiny
🎯 SPAR-Bench-Tiny
A lightweight subset of SPAR-Bench for fast evaluation of spatial reasoning in vision-language models (VLMs).
SPAR-Bench-Tiny contains 1,000 manually verified QA pairs — 50 samples per task across 20 spatial tasks — covering single-view and multi-view inputs.
This dataset mirrors the structure and annotation of the full SPAR-Bench, but is 10× smaller, making it ideal for low-latency evaluation.
📥 Load with… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-Bench-Tiny.tiny-imagenet
Dataset Card for tiny-imagenet
Dataset Summary
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images. Each class has 500 training images, 50 validation images, and 50 test images.
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=64x64 at 0x1A800E8E190,
'label': 15
}… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tiny-imagenet.conceptual_captions_3m_zh_tinyplantvillage-tiny
PlantVillage (tiny)
This is a debug-grade subset, not a faithful subsample for analysis.
50 images per class drawn from the full PlantVillage dataset is too few
to represent class-level visual diversity. Use it for iterating on
training-loop code, smoke-testing pipelines, or any situation where you
want the data structure but not the data scale. For actual classifier
training or evaluation, use
geraldmc/plantvillage-full.
What's in this dataset
A stratified subsample… See the full description on the dataset page: https://huggingface.co/datasets/geraldmc/plantvillage-tiny.OpenCOCO-I2T-Repack-Tiny
OpenCOCO-I2T-Repack-Tiny
OpenCOCO-I2T-Repack-Tiny is a compact image-to-text / image-text-to-text captioning dataset containing 158,958 image samples sourced from the COCO dataset and repackaged into a lightweight format suitable for vision-language model (VLM) fine-tuning. The dataset contains synthesized responses generated using a custom Qwen3.5 multimodal captioning pipeline. The input images undergo lossless image compression to significantly reduce the overall storage… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenCOCO-I2T-Repack-Tiny.TinyTown20x20qaconceptual_captions_3m_zh_tiny_0
Dataset Card for "conceptual_captions_3m_zh_tiny_0"
More Information needed
Caption3o-Opt-v3-Tiny
Caption3o-Opt-v3-Tiny
Caption3o-Opt-v3-Tiny is a high-quality, compact image-caption dataset designed for training and evaluating image-to-text models. Derived from prithivMLmods/blip3o-caption-mini-arrow and other curated sources, this optimized tiny version emphasizes long-form captions and covers a wide range of real-world and artistic scenes.
Dataset Summary
Size: 27,048 image-caption pairs
Format: Parquet
Image resolution: 512x512
Languages: English
Modality:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Caption3o-Opt-v3-Tiny.ling-3.0-tiny-atlas
ling-3.0-tiny-atlas
Molmo2-SynMultiImageQA-tinySFHQ-Tiny-512-Part1conceptual_captions_3m_zh_tiny_2
Dataset Card for "conceptual_captions_3m_zh_tiny_2"
More Information needed
conceptual_captions_3m_zh_tiny_5
Dataset Card for "conceptual_captions_3m_zh_tiny_5"
More Information needed
holyC-tinyllama-two-layer
HolyC TinyLlama Two-Layer Release
This bundle packages the HolyC TinyLlama work as a two-stage stack with the datasets that fed it. The goal is simple: make the release feel polished, uploadable, and honest about how it was built.
layer1/: explanatory adapter tuned for HolyC code understanding and explanation
layer2/: completion-oriented adapter tuned for HolyC code generation tasks
datasets/codebase/: raw HolyC code corpus
datasets/explanations/: explanation-oriented instruction… See the full description on the dataset page: https://huggingface.co/datasets/Aptlantis/holyC-tinyllama-two-layer.comix_v0_tiny_pages
Comic Books Tiny Dataset v0 - Pages (Testing)
Small test dataset of comic book pages for rapid development and testing.
⚠️ This is a TINY dataset for testing only. For production, use comix_v0_pages.
What's Included
Each page has:
{page_id}.jpg - Page image
{page_id}.json - Metadata (detections, captions, page class)
{page_id}.seg.npz - Segmentation masks (SAMv2)
Quick Start
from datasets import load_dataset
import numpy as np
# Load tiny pages dataset
pages… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix_v0_tiny_pages.tinytlpimagenet-1k_tiny
Dataset Card for "imagenet-1k_mini_100"
More Information needed
tiny-imagenet-200-clean
Dataset Card for tiny-imagenet-200-clean
Dataset Summary
The original Tiny ImageNet contained 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images. Each class has 500 training images, 50 validation images, and 50 test images.
This clean version removed grey scale images and only kept RGB images.
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
{
'image':… See the full description on the dataset page: https://huggingface.co/datasets/slegroux/tiny-imagenet-200-clean.SPAR-Bench-Tiny-RGBD
🎯 SPAR-Bench-Tiny-RGBD
A lightweight RGBD version of SPAR-Bench for fast evaluation of 3D-aware spatial reasoning in vision-language models (VLMs).
SPAR-Bench-Tiny-RGBD is a subset of SPAR-Bench-RGBD, containing 1,000 QA samples (50 per task × 20 tasks), each augmented with depths, camera intrinsics, and pose information.This dataset is ideal for quick evaluation of 3D-aware models, while maintaining compatibility with the same structure as… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-Bench-Tiny-RGBD.
