datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-imagenet
Dataset Card for tiny-imagenet
Dataset Summary
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images. Each class has 500 training images, 50 validation images, and 50 test images.
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=64x64 at 0x1A800E8E190,
'label': 15
}… See the full description on the dataset page: https://huggingface.co/datasets/zh-plus/tiny-imagenet.product-photography-v1-tiny-prompts-tasks-collage-filteredTiny-GenImage
Tiny GenImage Dataset
📝 Dataset Description
Dataset Summary
The Tiny GenImage Dataset is a curated, scaled-down collection of images and associated metadata designed to train, validate, and benchmark models for detecting and identifying artificially generated content. The dataset contains a mix of real-world images alongside those generated by prominent AI models, including various diffusion models (like Stable Diffusion 1.4/1.5, GLIDE, Midjourney, ADM, VQDM… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/Tiny-GenImage.text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
TinyMixtral-4x248M-MoE-atlas
juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas
A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing.
If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.tiny-audits
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DivyaApp/tiny-audits.conceptual_captions_3m_en_tinytiny-imagenet
Dataset Card for tiny-imagenet
Dataset Summary
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images. Each class has 500 training images, 50 validation images, and 50 test images.
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=64x64 at 0x1A800E8E190,
'label': 15
}… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tiny-imagenet.SPAR-Bench-Tiny
🎯 SPAR-Bench-Tiny
A lightweight subset of SPAR-Bench for fast evaluation of spatial reasoning in vision-language models (VLMs).
SPAR-Bench-Tiny contains 1,000 manually verified QA pairs — 50 samples per task across 20 spatial tasks — covering single-view and multi-view inputs.
This dataset mirrors the structure and annotation of the full SPAR-Bench, but is 10× smaller, making it ideal for low-latency evaluation.
📥 Load with… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-Bench-Tiny.conceptual_captions_3m_zh_tinyplantvillage-tiny
PlantVillage (tiny)
This is a debug-grade subset, not a faithful subsample for analysis.
50 images per class drawn from the full PlantVillage dataset is too few
to represent class-level visual diversity. Use it for iterating on
training-loop code, smoke-testing pipelines, or any situation where you
want the data structure but not the data scale. For actual classifier
training or evaluation, use
geraldmc/plantvillage-full.
What's in this dataset
A stratified subsample… See the full description on the dataset page: https://huggingface.co/datasets/geraldmc/plantvillage-tiny.OpenCOCO-I2T-Repack-Tiny
OpenCOCO-I2T-Repack-Tiny
OpenCOCO-I2T-Repack-Tiny is a compact image-to-text / image-text-to-text captioning dataset containing 158,958 image samples sourced from the COCO dataset and repackaged into a lightweight format suitable for vision-language model (VLM) fine-tuning. The dataset contains synthesized responses generated using a custom Qwen3.5 multimodal captioning pipeline. The input images undergo lossless image compression to significantly reduce the overall storage… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenCOCO-I2T-Repack-Tiny.conceptual_captions_3m_zh_tiny_0
Dataset Card for "conceptual_captions_3m_zh_tiny_0"
More Information needed
Caption3o-Opt-v3-Tiny
Caption3o-Opt-v3-Tiny
Caption3o-Opt-v3-Tiny is a high-quality, compact image-caption dataset designed for training and evaluating image-to-text models. Derived from prithivMLmods/blip3o-caption-mini-arrow and other curated sources, this optimized tiny version emphasizes long-form captions and covers a wide range of real-world and artistic scenes.
Dataset Summary
Size: 27,048 image-caption pairs
Format: Parquet
Image resolution: 512x512
Languages: English
Modality:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Caption3o-Opt-v3-Tiny.Molmo2-SynMultiImageQA-tinyling-3.0-tiny-atlas
ling-3.0-tiny-atlas
conceptual_captions_3m_zh_tiny_2
Dataset Card for "conceptual_captions_3m_zh_tiny_2"
More Information needed
conceptual_captions_3m_zh_tiny_5
Dataset Card for "conceptual_captions_3m_zh_tiny_5"
More Information needed
tiny-imagenet-200-clean
Dataset Card for tiny-imagenet-200-clean
Dataset Summary
The original Tiny ImageNet contained 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images. Each class has 500 training images, 50 validation images, and 50 test images.
This clean version removed grey scale images and only kept RGB images.
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
{
'image':… See the full description on the dataset page: https://huggingface.co/datasets/slegroux/tiny-imagenet-200-clean.imagenet-1k_tiny
Dataset Card for "imagenet-1k_mini_100"
More Information needed
SPAR-Bench-Tiny-RGBD
🎯 SPAR-Bench-Tiny-RGBD
A lightweight RGBD version of SPAR-Bench for fast evaluation of 3D-aware spatial reasoning in vision-language models (VLMs).
SPAR-Bench-Tiny-RGBD is a subset of SPAR-Bench-RGBD, containing 1,000 QA samples (50 per task × 20 tasks), each augmented with depths, camera intrinsics, and pose information.This dataset is ideal for quick evaluation of 3D-aware models, while maintaining compatibility with the same structure as… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhango/SPAR-Bench-Tiny-RGBD.ade20k-tiny
Dataset Card for ADE 20K Tiny
This is a tiny subset of the ADE 20K dataset, which you can find here.
test-finepersonas-v0.1-tiny-flux-schnell
Dataset Card for test-finepersonas-v0.1-tiny-flux-schnell
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/plaguss/test-finepersonas-v0.1-tiny-flux-schnell/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/test-finepersonas-v0.1-tiny-flux-schnell.finepersonas-v0.1-tiny-flux-schnell
Dataset Card for finepersonas-v0.1-tiny-flux-schnell
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds =… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/finepersonas-v0.1-tiny-flux-schnell.tiny-imagenet
Dataset Card for tiny-imagenet
Dataset Summary
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images. Each class has 500 training images, 50 validation images, and 50 test images.
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=64x64 at 0x1A800E8E190,
'label': 15
}… See the full description on the dataset page: https://huggingface.co/datasets/thethinkmachine/tiny-imagenet.VGBDataset-Tiny-HF[Paper] - [Website]
DocLayNet-tiny
Dataset Card for "DocLayNet-tiny"
Tiny set for unit tests based on https://huggingface.co/datasets/pierreguillou/DocLayNet-small.
Total ~0.1% of DocLayNet.
tinyworlds_parquet
TinyWorlds (parquet)
Retro game frames from AlmondGod/tinyworlds,
repackaged into a uniform per-frame parquet schema with one split per game for fast,
random-access frame loading.
TinyWorlds is a Genie reimplementation: it has no action
labels and learns latent actions from video alone. These splits are therefore action-less --
ideal for training an image tokenizer (RAE/VAE) or an unconditional video world model.
Splits
split
resolution
~fps
actions… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/tinyworlds_parquet.tiny-ai2d-irttiny-gqa-bounding-box-detailed
Tiny GQA Bounding Box (Detailed)
Dataset Summary
Tiny GQA Bounding Box (Detailed) is a subset of the GQA dataset enriched with object-level grounding information. Each example pairs a question-answer instance with a key object, including its bounding box and semantic label, derived from scene graphs.
This dataset is designed for research in:
Visual Question Answering (VQA)
Multimodal reasoning
Grounded reasoning and interpretability
Object-centric evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/Oztobuzz/tiny-gqa-bounding-box-detailed.
