datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.laion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.Holistic-Processing-Illusion-Faces
HoloFaceIllusion-Bench-EEG
A large-scale benchmark of holistic-face illusion stimuli for testing
human-vs-DNN alignment on configural face processing and for paired
EEG-decoder evaluation. Built entirely with classical CV
(dlib landmarks + MediaPipe Face Mesh + InsightFace gender/age +
OpenCV Poisson cloning + Reinhard LAB colour transfer) — no neural
networks or generative AI are used to create any pixel.
Three paradigms are included:
Paradigm
Cases
Conditions per case… See the full description on the dataset page: https://huggingface.co/datasets/Enhui-1/Holistic-Processing-Illusion-Faces.cifar100-enrichedThe CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).oxford-pets
Oxford-IIIT Pet Dataset
Images from The Oxford-IIIT Pet Dataset. Only images and labels have been pushed, segmentation annotations were ignored.
Homepage: https://www.robots.ox.ac.uk/~vgg/data/pets/
License:
Same as the original dataset.
mini_imagenet
Dataset Card for "mini_imagenet"
More Information needed
english-manhwa-scrape
English Manhwa Scrape Dataset
Request More ScrapesOrder Private Scrapes
Note: The total files are about 150+GB and my internet connection is slow af. So I'll be uploading them in batches
File Structure:
files/
- Manhwa 1 Name
- Chapter 1.cbz
- Chapter 2.cbz
- Chapter 3.cbz
- ......
- Chapter n.cbz
- Manhwa 2 Name
- Chapter 1.cbz
- Chapter 2.cbz
- Chapter 3.cbz
- ......
- Chapter n.cbz
Due to maximum 10000 files in a repo limit of huggingface, I had to further… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/english-manhwa-scrape.oxford-iiit-pet-vl-enriched
Visualize on Visual Layer
Oxford-IIIT-Pets-VL-Enriched
An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues!
With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.English-Handwritten-Math-Notes-Dataset
English Handwritten Math Notes Dataset
This dataset contains high-resolution images of handwritten mathematical notes written in English. It includes problem statements, worked examples, formulas, and annotated derivations. The dataset supports AI research in handwriting recognition, mathematical OCR, and document understanding for STEM applications.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/English-Handwritten-Math-Notes-Dataset.1-engineering-repair-50-commercial
Industrial Mechanical & Electrical Components Dataset — 50 Authentic Details
📋 Description
Curated collection of 50 high-resolution photographs documenting authentic industrial mechanical and electrical components — from electric motors and solenoid valves to circuit boards, wiring harnesses, ball bearings, gearboxes, caster wheels, and battery packs.
Each image includes comprehensive CSV metadata with 13 classification fields optimized for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Kos1976/1-engineering-repair-50-commercial.cifar10-enrichedThe CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images
per class. There are 50000 training images and 10000 test images.
This version if CIFAR-10 is enriched with several metadata such as embeddings, baseline results and label error scores.engineering-drawings-as1100
Engineering Drawings AS1100 Compliance Dataset
Dataset Description
This dataset contains engineering drawings with various AS1100 (Australian Standard for Technical Drawing) compliance issues for training AI models to identify missing elements and non-compliance issues in technical drawings.
Dataset Summary
The Engineering Drawings AS1100 Compliance Dataset is designed to train and evaluate vision-language models on identifying compliance issues in technical… See the full description on the dataset page: https://huggingface.co/datasets/jcrzd/engineering-drawings-as1100.fer2013-enhanced
FER2013 Enhanced: Advanced Facial Expression Recognition Dataset
The most comprehensive and quality-enhanced version of the famous FER2013 dataset for state-of-the-art emotion recognition research and applications.
🎯 Dataset Overview
FER2013 Enhanced is a significantly improved version of the landmark FER2013 facial expression recognition dataset. This enhanced version provides AI-powered quality assessment, balanced data splits, comprehensive metadata, and multi-format… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/fer2013-enhanced.GLAMI-Entity-Matching-Dataset
GLAMI Duplication Detection
Product-duplicate detection over GLAMI e-commerce listings: ~1.3M product
images plus multilingual titles, descriptions and attributes, with labelled
groups of items that do or do not refer to the same physical product.
Released under the Apache License 2.0 — see LICENSE.
TODO: describe how the labels were produced.
Structure
Config
Files
Contents
images
images/shard-*.parquet
itemId → image bytes, one row per product image… See the full description on the dataset page: https://huggingface.co/datasets/zidcenek/GLAMI-Entity-Matching-Dataset.bilingual-ocr-ru-en-synthetic
Bilingual OCR RU-EN Synthetic Dataset
This synthetic dataset is designed for bilingual text recognition (OCR) and script classification tasks (cyrillic / latin) at the word and short-line level.
Why are numbers, mathematical symbols, and the Greek alphabet included in the generation?
When creating synthetic OCR datasets, including an expanded set of characters (digits, mathematical signs, and Greek letters) is a deliberate step aimed at two main goals:… See the full description on the dataset page: https://huggingface.co/datasets/tehnik-tehnolog/bilingual-ocr-ru-en-synthetic.3-engineering-repair-51-commercial
Industrial Mechanical & Electrical Components Dataset — 51 Authentic Details
📋 Dataset Summary
Curated collection of 51 high-resolution photographs documenting authentic industrial mechanical and electrical components — electric motors and rotors, ball bearings and pulleys, gears and gearboxes, solenoid valves and hydraulic fittings, circuit boards and wiring harnesses, fasteners and tools, timing belts and caster wheels, air filters and brush assemblies.
Each… See the full description on the dataset page: https://huggingface.co/datasets/HerryChenwang111/3-engineering-repair-51-commercial.autotrain-data-enchondroma-vs-low-grade-chondrosarcoma-histology
AutoTrain Dataset for project: enchondroma-vs-low-grade-chondrosarcoma-histology
Dataset Description
This dataset has been automatically processed by AutoTrain for project enchondroma-vs-low-grade-chondrosarcoma-histology.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<1024x1024 RGB PIL image>",
"target": 0
},
{… See the full description on the dataset page: https://huggingface.co/datasets/itslogannye/autotrain-data-enchondroma-vs-low-grade-chondrosarcoma-histology.en-cldi
Dataset Description
"EN-CLDI" (en-cldi) contains 1690 classes, which contains images paired with verbs and adjectives. Each word within this set is unique and paired with at least 22 images.
It is the English subset of CLDI (cross-lingual dictionary induction) dataset from (Hartmann and Søgaard, 2018).
How to Use
from datasets import load_dataset
# Load the dataset
common_words = load_dataset("jaagli/en-cldi", split="train")
Citation… See the full description on the dataset page: https://huggingface.co/datasets/jaagli/en-cldi.LeafNetThe PlantVillage dataset, with over 54,000 images spanning 14 plant species and 26 disease types, has been widely used for leaf disease classification. However, it is limited in both scale and diversity. To address these limitations, we developed LeafNet, a large-scale dataset designed to support foundation models for leaf disease diagnosis. We introduce LeafNet comprises over 186,000 images from 22 crop species, covering 43 fungal diseases, 8 bacterial diseases, 2 mould (oomycete) diseases, 6… See the full description on the dataset page: https://huggingface.co/datasets/enalis/LeafNet.entitynet
EntityNet: Using Knowledge Graphs to harvest datasets for efficient CLIP model training
Dataloader and instructions can be found here: https://github.com/lmb-freiburg/entitynet
If you use the dataset, code, or results, please cite:
@inproceedings{ging2025entitynet,
author = {Simon Ging and Sebastian Walter and Jelena Bratuli{\'c} and Johannes Dienert and Hannah Bast and Thomas Brox},
title = {Using Knowledge Graphs to Harvest Datasets for Efficient {CLIP} Model… See the full description on the dataset page: https://huggingface.co/datasets/lmb-freiburg/entitynet.encyclopaedia_britannica_illustrated
Dataset Card: Encyclopaedia Britannica Illustrated
Homepage: https://data.nls.uk/data/digitised-collections/encyclopaedia-britannica/
Dataset Summary
A binary image classification dataset containing 2,573 images from the National Library of Scotland's Encyclopaedia Britannica collection (1768-1860). Each image is labeled as either "illustrated" or "not-illustrated". The dataset was created for experimenting with historical document image classification.… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/encyclopaedia_britannica_illustrated.3-engineering-repair-51-commercial
Industrial Mechanical & Electrical Components Dataset — 51 Authentic Details
📋 Dataset Summary
Curated collection of 51 high-resolution photographs documenting authentic industrial mechanical and electrical components — electric motors and rotors, ball bearings and pulleys, gears and gearboxes, solenoid valves and hydraulic fittings, circuit boards and wiring harnesses, fasteners and tools, timing belts and caster wheels, air filters and brush assemblies.
Each… See the full description on the dataset page: https://huggingface.co/datasets/Kos1976/3-engineering-repair-51-commercial.chars74k-eng-good
Chars74k
The "Good" subset of the "English" subset of the Chars74k dataset split into training and validation sets.
The validation set was created to match the label distribution of the training set.
62 classes (0-9, A-Z, a-z)
Dataset page: https://teodecampos.github.io/chars74k/
Paper describing the dataset:https://www.semanticscholar.org/paper/Character-Recognition-in-Natural-Images-Campos-Babu/dbbd5fdc09349bbfdee7aa7365a9d37716852b32
5 images where removed due to poor quality.… See the full description on the dataset page: https://huggingface.co/datasets/Lajdre/chars74k-eng-good.enhanced-indian-food-classification
Enhanced Indian Food Classification Dataset
A comprehensive dataset for Indian food classification with 15,404 images across 43 classes.
Dataset Structure
dataset/
├── train/ # Training images
├── validation/ # Validation images
└── test/ # Test images
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("SohlHealth/enhanced-indian-food-classification")
# Access splits
train_data = dataset['train']
val_data… See the full description on the dataset page: https://huggingface.co/datasets/SohlHealth/enhanced-indian-food-classification.food101-enriched
Dataset Card for Food-101-Enriched (Enhanced by Renumics)
Dataset Summary
📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML… See the full description on the dataset page: https://huggingface.co/datasets/renumics/food101-enriched.I-SPY_2_Breast_Dynamic_Contrast_Enhanced_MRI_Trial_T0_and_T3_DCE-MRI_Datasetengineering-drawings-as1100
Engineering Drawings AS1100 Compliance Dataset
Dataset Description
This dataset contains engineering drawings with various AS1100 (Australian Standard for Technical Drawing) compliance issues for training AI models to identify missing elements and non-compliance issues in technical drawings.
Dataset Summary
The Engineering Drawings AS1100 Compliance Dataset is designed to train and evaluate vision-language models on identifying compliance issues in… See the full description on the dataset page: https://huggingface.co/datasets/Vasanthaleela/engineering-drawings-as1100.Medical-Multimodal-EN-TH
HealthGPTVL-Translation Medical-Multimodal-EN-TH
This dataset is a bilingual (English-Thai) medical multimodal evaluation dataset containing medical images with corresponding question-answer pairs for visual question answering and translation tasks.
Dataset Details
Dataset Description
This dataset contains 17,047 medical image-text pairs designed for multimodal medical AI evaluation. It includes medical images from various imaging modalities (MRI, CT, X-Ray… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-Multimodal-EN-TH.16-Amadeus-Performance-Art-65-Enterprise-CLEAR-Compliant-POC
🌫️ Privacy-Native Blurred Environments: Compliance-Native Multimodal Tokens (POC)
🛡️ Engineering Evaluation Sandbox (Active 7-Day Access)
Technical Ingestion Portal: s3://createphotos (Whitelisted buckets only)
Secure Evaluation Link: Download Privacy_Native_Blurred_Environments_POC.zip
Direct Manifest Auditor: BWS Forensic Manifest Repository
Procurement: All assets are 2026 US CLEAR Act compliant. Access is granted to whitelisted engineering nodes only. Forward your AWS… See the full description on the dataset page: https://huggingface.co/datasets/BWS-Data-Solutions/16-Amadeus-Performance-Art-65-Enterprise-CLEAR-Compliant-POC.
