datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mnist-cleaned-full
Dataset Card for 2025.11.21.16.40.44.970939
This is a FiftyOne dataset with 69807 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Linus-L/mnist-cleaned-full")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Linus-L/mnist-cleaned-full.line-ex
LineEX
This dataset repo contains:
the uploaded train split with 397,993 images
the released test split with 20,000 images
Repo: 13point5/line-ex
Shared Schema
image_id
file_name
image
width
height
data_type
chart_elements
lines
chart_elements fields
annotation_id
category_id
category_name
bbox_xywh
area
text
line_id
lines fields
annotation_id
category_id
category_name
line_name
polyline_xy
raw_series_xy
area
Notes
The repo… See the full description on the dataset page: https://huggingface.co/datasets/13point5/line-ex.Linnaeus5
Linnaeus 5
A dataset containing images in 5 categories: berry, bird, dog, flower, and other. Data taken from here
Dataset Structure
256x256 pixels images
5 classes: berry, bird, dog, flower, other
Train and test splits
2026.mechaptcha.linear-probe-experiments-giant-20260525
siddharthmb/2026.mechaptcha.linear-probe-experiments-giant-20260525
Paired CAPTCHA image experiments for linear probes over a trained Mechaptcha CNN.
Each example contains a matched image_a and image_b pair generated from the same seed pool.
Intended Use
This dataset is designed for linear probe experiments that compare activations from Batch A against Batch B.
Use label 1 for image_a and label 0 for image_b.
Recommended checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.mechaptcha.linear-probe-experiments-giant-20260525.mnist-cleaned-up
Dataset Card for cleaned-up-mnist-training-set
This is a FiftyOne dataset with 505 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Linus-L/mnist-cleaned-up")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Linus-L/mnist-cleaned-up.autotrain-data-multifamily_v2
AutoTrain Dataset for project: multifamily_v2
Dataset Description
This dataset has been automatically processed by AutoTrain for project multifamily_v2.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<500x333 RGB PIL image>",
"target": 33
},
{
"image": "<500x667 RGB PIL image>",
"target": 11
}]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lineups-io/autotrain-data-multifamily_v2.CHUBS
CHUBS: A Large-Scale Dataset of Chu Bamboo Slip Script
Code | Paper (upcoming)
Introduction
This is a large-scale dataset of Chu bamboo slip (CBS, Chinese: 楚简, chujian) script, an ancient Chinese script used during the Spring and Autumn period over 2,000 years ago. This dataset consists of two parts:
The main dataset where each example is an image and the corresponding text label. This part is contained in the glyphs.zip ZIP file.
A character detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/LINGJIAN5028/CHUBS.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.NWPU-RESISC45
Dataset Card for "NWPU-RESISC45"
Licensing Information
[CC-BY-SA]
Citation Information
Remote sensing image scene classification: Benchmark and state of the art
@article{cheng2017remote,
title = {Remote sensing image scene classification: Benchmark and state of the art},
author = {Cheng, Gong and Han, Junwei and Lu, Xiaoqiang},
year = 2017,
journal = {Proceedings of the IEEE},
publisher = {IEEE},
volume… See the full description on the dataset page: https://huggingface.co/datasets/Ling200424/NWPU-RESISC45.sportwear_lineart_dataset
Sportwears Mates Lineart Dataset
Now this is a different type of image dataset.
Here it's not beautiful images, nor taggings.
These are spreadsheets in multiple views of sportsman in sportswear.
Oh, I admit this soundtrack sounds epic now (non-commercial, not part of license, no derivative work allowed)
Title: Stand As One (Stadium Chant Edition)
Despite not being exactly good for image generative models (which generate images), the idea is the other end of such chain.
It… See the full description on the dataset page: https://huggingface.co/datasets/eastenddan/sportwear_lineart_dataset.lineex-test
LineEX Test Split
This dataset is a structured Hugging Face mirror of the test split released with the
LineEX paper, "LineEX: Data Extraction from Scientific Line Charts" (WACV 2023).
Repo: 13point5/lineex-test
Split: test
Rows: 20,000
Source format: PNG images plus COCO-style chart-element and line annotations
Columns
image_id
file_name
image
width
height
data_type
chart_elements
lines
chart_elements fields
annotation_id
category_id
category_name… See the full description on the dataset page: https://huggingface.co/datasets/13point5/lineex-test.Technical-Textile-Synthesis-Luxury-Lingerie-Embroidery
🩰 BWS La Perla Premium: Compliance-Native Multimodal Tokens (POC)
🛡️ Engineering Evaluation Sandbox (Active 7-Day Access)
Technical Ingestion Portal: s3://createphotos (Whitelisted buckets only)
Secure Evaluation Link: Download 03_La_Perla__lingerie_30_Enterprise_POC.zip
Direct Manifest Auditor: BWS Forensic Manifest Repository
Procurement: All assets are 2026 US CLEAR Act compliant. Access is granted to whitelisted engineering nodes only. Forward your AWS Account ID to… See the full description on the dataset page: https://huggingface.co/datasets/BWS-Data-Solutions/Technical-Textile-Synthesis-Luxury-Lingerie-Embroidery.first-set-RMBThis dataset includes images of genuine and counterfeit currency from the first set of RMB/Renminbi. Part of the genuine currency are samples of prototype banknotes.
The first set of RMB/Renminbi is incredibly rare, and gathering data has been challenging. Your contributions are welcome.
autotrain-data-multifamily
AutoTrain Dataset for project: multifamily
Dataset Description
This dataset has been automatically processed by AutoTrain for project multifamily.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<500x500 RGB PIL image>",
"target": 40
},
{
"image": "<500x500 RGB PIL image>",
"target": 34
}]
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/lineups-io/autotrain-data-multifamily.diffusion_model_assessment_v2
Dataset Card for generated_flowers_with_embeddings
This is a FiftyOne dataset with 20 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Linus-L/diffusion_model_assessment_v2")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Linus-L/diffusion_model_assessment_v2.resisc45
Description
RESISC45 dataset is a publicly available benchmark for Remote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This dataset contains 31,500 images, covering 45 scene classes with 700 images in each class.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/Lin1019/resisc45.DisasterM3
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
Junjue Wang*,
Weihao Xuan*,
Heli Qi, Zhihao Liu, Kunyi Liu, Yuhan Wu, Hongruixuan Chen,
Jian Song
Junshi Xia, Zhuo Zheng, Naoto Yokoya†
* Equal Contributions
† Corresponding Author
Paper: https://arxiv.org/abs/2505.21089
Code: https://github.com/Junjue-Wang/DisasterM3
Highlights
DisasterM3 includes 26,988 bi-temporal satellite images and 123k instruction pairs across… See the full description on the dataset page: https://huggingface.co/datasets/Ling200424/DisasterM3.one-class-eus
One-Class EUS
This dataset contains de-identified endoscopic ultrasound images for medical one-class classification experiments.
Contents
images/: 1,450 resized PNG images.
metadata.csv: full image list with labels.
dataset_summary.json: dataset counts.
Labels
Label
Count
non_gist_lesion
885
gist
565
Metadata Format
metadata.csv contains:
Column
Description
file_name
Relative image path for HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/linfan150/one-class-eus.
