datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
konachan_full
konachan Full Dataset
This is the full dataset of konachan.com. And all the original images are maintained here.
Information
Images
There are 319040 images in total. The maximum ID of these images is 391069. Last updated at 2025-07-21 22:19:40 JST.
These are the information of recent 50 images:
id
filename
width
height
mimetype
tags
file_url
391069
391069.png
4774
2786
image/png
animal anthropomorphism azur_lane bird black_hair bondage building car… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/konachan_full.open-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.sankaku_full
Sankaku Full Dataset
This is the full dataset of chan.sankakucomplex.com. And all the original images are maintained here.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into HF_TOKEN environment variable to give the code authorization for this repository.
from cheesechaser.datapool import SankakuDataPool… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/sankaku_full.zerochan_full
Zerochan Full Dataset
This is the full dataset of zerochan.net. And all the original images are maintained here.
Thanks to @AngelBottomless for supporting the computational resources for the creation of this dataset.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into HF_TOKEN environment variable to give… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/zerochan_full.gelbooru_full
Gelbooru Full Dataset
This is the full dataset of gelbooru.com. And all the original images are maintained here.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into HF_TOKEN environment variable to give the code authorization for this repository.
from cheesechaser.datapool import GelbooruDataPool
pool =… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/gelbooru_full.plantvillage-full
PlantVillage (full)
A curated re-release of the PlantVillage plant-disease image dataset
(Mohanty, Hughes, Salathé 2016), with structured per-image metadata and
a leaf-grouped train/test split. Built for the iResearch Institute 2026
Virtual Lab mentorship's Rationale 2 track. The companion debug-grade
subset is at
geraldmc/plantvillage-tiny.
What's in this dataset
54,304 images of plant leaves, photographed against plain backgrounds
under controlled lighting, across… See the full description on the dataset page: https://huggingface.co/datasets/geraldmc/plantvillage-full.mnist-cleaned-full
Dataset Card for 2025.11.21.16.40.44.970939
This is a FiftyOne dataset with 69807 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Linus-L/mnist-cleaned-full")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Linus-L/mnist-cleaned-full.pexels-tagger-v0-w640-ws-full
Pexels Tagger V0 Webdataset Full Dataset
This is the webdataset dataset for animetimm/pexels-wdtagger-w640.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('animetimm/pexels-tagger-v0-w640-ws-full')
print(dataset["train"][0])
Images
3122908 images in total.
Split
Image Count
Total Size
train
2810634
184 GB
test
156409
10.2 GB
val
155865
10.2 GB
Tags… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/pexels-tagger-v0-w640-ws-full.rule34_full
Rule34 Full Dataset
This is the full dataset of rule34.xxx. And all the original images are maintained here.
Information
Images
There are 11336807 images in total. The maximum ID of these images is 13078768. Last updated at 2025-04-10 21:23:24 JST.
These are the information of recent 50 images:
id
filename
width
height
mimetype
tags
file_size
file_url
13078768
13078768.jpeg
1024
1024
image/jpeg
1boy 1girls ai_generated ass bubble_butt… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/rule34_full.vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
pd-voice-full-multimodal-dataset
Parkinson Voice — Full Multimodal Dataset
Complete Parkinson’s vs healthy voice package for classification and explainable Mel reasoning research (EDGE).
Not Mel-only: raw audio, 10 visual modalities, feature CSVs, plus Gemma reasoning traces for Mel.
Contents
Path
Description
audio/
Waveform clips (healthy / parkinsons), 1134 files
images/mel/
Mel spectrograms
images/spectrogram/
Linear spectrograms
images/mfcc/
MFCC maps
images/delta_mfcc/… See the full description on the dataset page: https://huggingface.co/datasets/mdimamhosen/pd-voice-full-multimodal-dataset.yande_full
Yande Full Dataset
This is the full dataset of yande.re. And all the original images are maintained here.
Information
Images
There are 1124996 images in total. The maximum ID of these images is 1236881. Last updated at 2025-07-25 04:05:43 JST.
These are the information of recent 50 images:
id
filename
width
height
mimetype
tags
file_url
1236881
1236881.jpg
1964
3508
image/jpeg
animal_ears bikini cloba nekomimi nipples see_through swimsuits tail… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/yande_full.plantdoc-full
PlantDoc — full variant
A curated mirror of the PlantDoc plant disease classification dataset (Singh et al. 2020), with normalized class metadata and a stable schema, hosted as a Hugging Face Dataset for reproducible distribution. The companion plantdoc-tiny variant is a 164-image stratified subsample of this dataset for fast test-suite use.
PlantDoc was built to address a specific failure mode in earlier plant disease datasets like PlantVillage: lab-condition images don't predict… See the full description on the dataset page: https://huggingface.co/datasets/geraldmc/plantdoc-full.house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.anime_pictures_full
Anime-Pictures Full Dataset
This is the full dataset of anime-pictures.net. And all the original images are maintained here.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into HF_TOKEN environment variable to give the code authorization for this repository.
from cheesechaser.datapool import… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/anime_pictures_full.e6ai_full
E6ai Dataset Newest Archive
This is the newest dataset of e6ai.net. And the newest data are up-to-date-ly maintained here, to make sure you can get all the newest data from huggingface instead of e6ai site.
All the file types we kept in this dataset: gif, jpg, png, webm
Information
Data
There are 110875 records in total. The ID range of these data is 3-126075. Last updated at 2025-09-09 01:08:59 JST.
These are the information of recent 50 records:
id… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/e6ai_full.google-illustrations-full
Google Illustrations: Complete 4K Multi-Layer Archive (2,052 Avatars)
A comprehensive, uncompressed 4096x4096 archival dataset of the complete Google Account Illustrations library, featuring all 59 artist collections, 2,052 unique scenes, decomposed layer assets, color presets, and semantic metadata.
Legal Disclaimer and Copyright Notice
PLEASE READ CAREFULLY:
This repository is an independent archival and educational compilation provided strictly for… See the full description on the dataset page: https://huggingface.co/datasets/Anson10124/google-illustrations-full.Full-Danbooru-Complement
Full Danbooru Complement
发布状态 / Release status(2026-08-06):基线版本、2026-06 月度归档和 2026-07 月度归档均已完成整理、验证并发布。
The baseline, 2026-06 monthly archive, and 2026-07 monthly archive have all been assembled, verified, and published.
概述 / Overview
Full Danbooru Complement 是 deepghs/danbooru2024-webp-4Mpixel 的后续补充数据集,用于延伸其 Danbooru 图像与 post metadata 覆盖范围。数据按不可变发布目录组织;每个 Parquet 行对应一个 Danbooru post,并在行内保存图像字节与相关 metadata。
Full Danbooru Complement is a follow-up complement to… See the full description on the dataset page: https://huggingface.co/datasets/Xuness/Full-Danbooru-Complement.safebooru_full
Safebooru Full Dataset
This is the full dataset of safebooru.org. And all the original images are maintained here.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into HF_TOKEN environment variable to give the code authorization for this repository.
from cheesechaser.datapool import SafebooruDataPool
pool =… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/safebooru_full.text-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.full-scene
DCSkyCam Full-Scene Wide-Angle Dataset
This dataset contains full-frame wide-angle images captured by the DCSkyCam — a Raspberry Pi-based webcam in Washington, DC. The images show the full camera field of view as captured by the HQ Camera module with a wide-angle lens.
Dataset Description
The DCSkyCam system used a three-stage detection pipeline:
Object Detection (SSD MobileNet V3) identifies candidate objects in the sky
Binary Classifier determines if a… See the full description on the dataset page: https://huggingface.co/datasets/dcskycam/full-scene.anima-style-embedding-500k-full-face
Anima Style Embedding Full-Frame and Face Corpus
Dataset summary
The corpus contains 500,000 person-bearing illustrations for open-set anime style-embedding research. Human artwork and Anima-generated images use independent style identities because an artist's original work and Anima's response to the corresponding artist tag are not the same visual distribution.
Source directory
Style identities
Images per identity
Images
synthetic/
5,000
50
250,000… See the full description on the dataset page: https://huggingface.co/datasets/ij/anima-style-embedding-500k-full-face.danbooru-wdtagger-v4-w640-ws-full
Danbooru WDTagger V4 Webdataset Full Dataset
This is the webdataset dataset for animetimm/danbooru-wdtagger-v4-w640.
Metadata here are cleaned by @SmilingWolf, can be used for multi-label classification training.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('animetimm/danbooru-wdtagger-v4-w640-ws-full')
print(dataset["train"][0])
Images
5914596 images in total.
Split… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/danbooru-wdtagger-v4-w640-ws-full.e621-wdtagger-v1-w640-ws-full
E621 WDTagger V1 Webdataset Full Dataset
This is the webdataset dataset for animetimm/e621-wdtagger-v1-w640.
Metadata here are cleaned by @SmilingWolf, can be used for multi-label classification training.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('animetimm/e621-wdtagger-v1-w640-ws-full')
print(dataset["train"][0])
Images
3714685 images in total.
Split
Image Count… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/e621-wdtagger-v1-w640-ws-full.house_kg_full_dataset_frames
house.kg — Kyrgyzstan Real Estate, over time
Sale and rental listings scraped from house.kg, the largest
real-estate board in Kyrgyzstan, re-measured on a schedule. Field names are
English; values are kept in the original language (Russian), exactly as the site
renders them.
Coverage: 2026-09-08. This is the baseline snapshot; later runs append new partitions.
Subsets
subset
rows
description
listings
25,264
one row per advertisement — current state plus… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset_frames.nozomi_standalone_full
Nozomi Full Dataset
This is the full dataset of nozomi.la. And only the standalone original images are maintained here.
Information
Images
There are 22435997 images in total. The maximum ID of these images is 36722740. Last updated at 2025-09-12 20:11:12 JST.
These are the information of recent 50 images:
id
filename
width
height
mimetype
tags
file_size
file_url
created_at
36722740
36722740.webp
2000
1363
image/webp
['1girl', 'all_fours'… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/nozomi_standalone_full.smdg-full-dataset
Dataset Card for Dataset Name
All the images of the dataset come from this kaggle dataset.
Only fundus images have been collected and some minor modifications have been made to the metadata.
All credit goes to the original authors and the contributor on Kaggle.
Dataset Details
Dataset Description
Standardized Multi-Channel Dataset for Glaucoma (SMDG-19) is a collection and standardization of 19 public datasets, comprised of full-fundus glaucoma images and… See the full description on the dataset page: https://huggingface.co/datasets/bumbledeep/smdg-full-dataset.fancaps_full
Fancaps Full Dataset
This is the full dataset of Fancaps. And all the original images are maintained here.
Information
Images
There are 29004217 images in total. The maximum ID of these images is 31449533. Last updated at 2025-09-05 08:51:12 JST.
These are the information of recent 50 images:
id
filename
width
height
type
rating
tags
has_face
face_count
face_width
face_height
face_min
face_max
31449533
31449533.jpg
1920
1080
image/jpeg
general… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/fancaps_full.TFQ-Data-Full
TFQ-Data: A Fine-Grained Dataset for Image Implication
TFQ-Data is a large-scale visual instruction tuning dataset specifically designed to train Multi-modal Large Language Models (MLLMs) on Image Implication and Metaphorical Reasoning.
Unlike standard VQA datasets that focus on literal description, TFQ-Data utilizes a True-False Question (TFQ) format. This format provides high knowledge density and verifiable reward signals, making it an ideal substrate for Visual Reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/MING-ZCH/TFQ-Data-Full.TFQ-Bench-Full
TFQ-Bench: A Benchmark for Evaluating Image Implication Understanding
TFQ-Bench is a rigorous evaluation benchmark designed to assess the capabilities of MLLMs in understanding visual metaphors, sarcasm, and implicit meanings via True-False Questions.
It serves as a complement to existing benchmarks like II-Bench (Multiple-Choice Question) and CII-Bench (Open-Style Question), offering a lower-bound difficulty check that tests a model's ability to verify specific propositions about… See the full description on the dataset page: https://huggingface.co/datasets/MING-ZCH/TFQ-Bench-Full.
