datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
plantvillage-full
PlantVillage (full)
A curated re-release of the PlantVillage plant-disease image dataset
(Mohanty, Hughes, Salathé 2016), with structured per-image metadata and
a leaf-grouped train/test split. Built for the iResearch Institute 2026
Virtual Lab mentorship's Rationale 2 track. The companion debug-grade
subset is at
geraldmc/plantvillage-tiny.
What's in this dataset
54,304 images of plant leaves, photographed against plain backgrounds
under controlled lighting, across… See the full description on the dataset page: https://huggingface.co/datasets/geraldmc/plantvillage-full.pexels-tagger-v0-w640-ws-full
Pexels Tagger V0 Webdataset Full Dataset
This is the webdataset dataset for animetimm/pexels-wdtagger-w640.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('animetimm/pexels-tagger-v0-w640-ws-full')
print(dataset["train"][0])
Images
3122908 images in total.
Split
Image Count
Total Size
train
2810634
184 GB
test
156409
10.2 GB
val
155865
10.2 GB
Tags… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/pexels-tagger-v0-w640-ws-full.plantdoc-full
PlantDoc — full variant
A curated mirror of the PlantDoc plant disease classification dataset (Singh et al. 2020), with normalized class metadata and a stable schema, hosted as a Hugging Face Dataset for reproducible distribution. The companion plantdoc-tiny variant is a 164-image stratified subsample of this dataset for fast test-suite use.
PlantDoc was built to address a specific failure mode in earlier plant disease datasets like PlantVillage: lab-condition images don't predict… See the full description on the dataset page: https://huggingface.co/datasets/geraldmc/plantdoc-full.house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.full-scene
DCSkyCam Full-Scene Wide-Angle Dataset
This dataset contains full-frame wide-angle images captured by the DCSkyCam — a Raspberry Pi-based webcam in Washington, DC. The images show the full camera field of view as captured by the HQ Camera module with a wide-angle lens.
Dataset Description
The DCSkyCam system used a three-stage detection pipeline:
Object Detection (SSD MobileNet V3) identifies candidate objects in the sky
Binary Classifier determines if a… See the full description on the dataset page: https://huggingface.co/datasets/dcskycam/full-scene.text-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.house_kg_full_dataset_frames
house.kg — Kyrgyzstan Real Estate, over time
Sale and rental listings scraped from house.kg, the largest
real-estate board in Kyrgyzstan, re-measured on a schedule. Field names are
English; values are kept in the original language (Russian), exactly as the site
renders them.
Coverage: 2026-09-08. This is the baseline snapshot; later runs append new partitions.
Subsets
subset
rows
description
listings
25,264
one row per advertisement — current state plus… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset_frames.smdg-full-dataset
Dataset Card for Dataset Name
All the images of the dataset come from this kaggle dataset.
Only fundus images have been collected and some minor modifications have been made to the metadata.
All credit goes to the original authors and the contributor on Kaggle.
Dataset Details
Dataset Description
Standardized Multi-Channel Dataset for Glaucoma (SMDG-19) is a collection and standardization of 19 public datasets, comprised of full-fundus glaucoma images and… See the full description on the dataset page: https://huggingface.co/datasets/bumbledeep/smdg-full-dataset.TFQ-Data-Full
TFQ-Data: A Fine-Grained Dataset for Image Implication
TFQ-Data is a large-scale visual instruction tuning dataset specifically designed to train Multi-modal Large Language Models (MLLMs) on Image Implication and Metaphorical Reasoning.
Unlike standard VQA datasets that focus on literal description, TFQ-Data utilizes a True-False Question (TFQ) format. This format provides high knowledge density and verifiable reward signals, making it an ideal substrate for Visual Reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/MING-ZCH/TFQ-Data-Full.fuliji
FuliJi (福利姬) Portrait Dataset
A dataset of Asian portrait photographs for fine-tuning Vision-Language Models.
Statistics
Metric
Value
Total Images
222
Unique Artists (福利姬)
222
Bilingual Descriptions
✅ EN + ZH
Schema
Column
Type
Description
image
image
Portrait photograph
fuliji
string
Artist name (福利姬)
gallery
string
Photo set/collection name
text_en
string
English description
text_zh
string
Chinese description… See the full description on the dataset page: https://huggingface.co/datasets/DownFlow/fuliji.TFQ-Bench-Full
TFQ-Bench: A Benchmark for Evaluating Image Implication Understanding
TFQ-Bench is a rigorous evaluation benchmark designed to assess the capabilities of MLLMs in understanding visual metaphors, sarcasm, and implicit meanings via True-False Questions.
It serves as a complement to existing benchmarks like II-Bench (Multiple-Choice Question) and CII-Bench (Open-Style Question), offering a lower-bound difficulty check that tests a model's ability to verify specific propositions about… See the full description on the dataset page: https://huggingface.co/datasets/MING-ZCH/TFQ-Bench-Full.reddit_mostlyhumans_full
Reddit Mostly-Humans Full Dataset
This is the full dataset of reddit_mostlyhumans dataset. And all the original images are maintained here.
Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous.
Information
Images
There are 3863451 images in total. The maximum ID of these images is 3863571. Last updated at 2024-10-26 17:18:31 UTC.
These are the information of recent 50 images:
id
filename
width
height
mimetype… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/reddit_mostlyhumans_full.midjourney_captioned_23m_full
Midjourney Captioned Full Dataset
This is the full dataset of Midjourney Captioned 23M dataset. And all the original images are maintained here.
Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous.
Information
Images
There are 23167456 images in total. The maximum ID of these images is 23167456. Last updated at 2024-12-01 12:11:43 UTC.
These are the information of recent 50 images:
id
width
height
filename… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.skin-lesion-trainOne-full
Skin Lesion Dataset — trainOne Full
All 7 HAM10000 classes including Melanocytic Nevi (nv).
Source: ISIC 2018 Challenge Task 3.
Code
Description
Count
nv
Melanocytic Nevi (benign)
5000
mel
Melanoma (deadly)
1113
bkl
Benign Keratosis
1099
bcc
Basal Cell Carcinoma (cancer)
514
akiec
Actinic Keratosis (monitor)
327
vasc
Vascular Lesions
142
df
Dermatofibroma
115
Total: 8,310 images
Load:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/zujiguts/skin-lesion-trainOne-full.cifar10-stats
CIFAR-10 CNN Layerwise Training Statistics
Dataset Description
This dataset contains layer-wise training statistics for a convolutional network trained on CIFAR-10, together with the corresponding test/acc.
Each row is one point at loss landscape. The features are computed on the last training batch of 1024 samples before the end of an epoch, and test/acc is measured immediately after that epoch.
The dataset includes statistics for several convolutional layers and the… See the full description on the dataset page: https://huggingface.co/datasets/Fullfix/cifar10-stats.
