datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndoLepAtlas
IndoLepAtlas — Indian Lepidoptera & Host Plants Dataset
A large-scale computer vision dataset of Indian butterflies, moths, and their larval host plants. Sourced from ifoundbutterflies.org with public CC-licensed photographs.
Inspired by: iNaturalist | Domain: Indian Wildlife & Biodiversity
Dataset Overview
Butterflies
Host Plants
Total
Species
961
127
1,088
Images
60,641
703
61,344
Source
ifoundbutterflies.org
ifoundbutterflies.org
—… See the full description on the dataset page: https://huggingface.co/datasets/Butterfree/IndoLepAtlas.indian-traditional-artificial-jewellery
Traditional and Handmade Indian Jewellery Dataset
This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks.
Dataset Overview
This dataset is designed to provide detailed information about Indian jewelry… See the full description on the dataset page: https://huggingface.co/datasets/Coder-Dragon/indian-traditional-artificial-jewellery.indoor-scene-classification
Dataset Labels
['meeting_room', 'cloister', 'stairscase', 'restaurant', 'hairsalon', 'children_room', 'dining_room', 'lobby', 'museum', 'laundromat', 'computerroom', 'grocerystore', 'hospitalroom', 'buffet', 'office', 'warehouse', 'garage', 'bookstore', 'florist', 'locker_room', 'inside_bus', 'subway', 'fastfood_restaurant', 'auditorium', 'studiomusic', 'airport_inside', 'pantry', 'restaurant_kitchen', 'casino', 'movietheater', 'kitchen', 'waitingroom', 'artstudio', 'toystore'… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/indoor-scene-classification.IIIT-INDIC-HW-WORDS-Hindi
IIIT-INDIC-HW-WORDS-Hindi
Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images.
Overview
The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research and… See the full description on the dataset page: https://huggingface.co/datasets/c3rl/IIIT-INDIC-HW-WORDS-Hindi.indian-traditional-artificial-jewellery
Traditional and Handmade Indian Jewellery Dataset
This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks.
Dataset Overview
This dataset is designed to provide detailed information about Indian… See the full description on the dataset page: https://huggingface.co/datasets/manidhardevu/indian-traditional-artificial-jewellery.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.INDOMEME
INDOMEME
INDOMEME is a multimodal dataset of Indonesian memes collected from Facebook, annotated for hate speech detection and content appropriateness classification. Each meme is enriched with OCR-extracted text and LLM-generated captions to support multimodal analysis.
Dataset Columns
Column
Description
image
Meme image
image_path
Original filename of the meme image
hate_final
Hatefulness label: hate or not hate
appropriate_final
Appropriateness… See the full description on the dataset page: https://huggingface.co/datasets/aiatums/INDOMEME.index-cards-harvard-botany-metropolitan-flora
Card File of the Flora of the Metropolitan Parks (Harvard Botany Libraries, 1894–1895)
4,574 botanical specimen index cards from the Harvard University Botany
Libraries' Card File of the Flora of the Metropolitan Parks, 1894–1895 (bulk),
compiled by Walter Deane (1848–1930). Records flora of the Metropolitan Park
system around Boston — Middlesex Fells Reservation, Blue Hills, Norfolk County,
and adjacent areas — with one card per specimen entry: species, locality,
collection date… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-harvard-botany-metropolitan-flora.gemini-2.5-vs-pro-realism-ai-perception
Gemini 2.5 Flash vs Gemini 3 Pro: Photorealism Comparaison
The dataset quantifies how much better Google's newest image model, Gemini 3 Pro Image, is at generating photorealistic images compared to Gemini 2.5 Flash Image.
This text-to-image benchmark dataset contains 6428 human judgments from annotators across 50+ countries, collected in under 20 minutes using the Rapidata Python API, accessible to anyone and ideal for small to large scale evaluation.
Overview
30… See the full description on the dataset page: https://huggingface.co/datasets/indomar/gemini-2.5-vs-pro-realism-ai-perception.enhanced-indian-food-classification
Enhanced Indian Food Classification Dataset
A comprehensive dataset for Indian food classification with 15,404 images across 43 classes.
Dataset Structure
dataset/
├── train/ # Training images
├── validation/ # Validation images
└── test/ # Test images
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("SohlHealth/enhanced-indian-food-classification")
# Access splits
train_data = dataset['train']
val_data… See the full description on the dataset page: https://huggingface.co/datasets/SohlHealth/enhanced-indian-food-classification.index-cards-parisian-parliamentarians
Parisian Parliamentarians — Scholarly Prosopography Index Cards
26 index cards (across 5 letter-range items A–D · E–H · J–O · P–R · S–Z) from
the parisianparliamentarians scholarly archive on the Internet Archive — a
prosopographical card index of Parisian parliamentary figures, originally compiled
as a research finding aid.
Each row pairs the full-resolution card scan with the Internet Archive's ABBYY OCR
text and full provenance back to the source IA item.
Source &… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-parisian-parliamentarians.index-cards-cas-galapagos-stewart-specimens
Alban Stewart's Galápagos Expedition Specimen Cards (CAS Archives, 1905–1906)
1,059 specimen index cards from the California Academy of Sciences Archives,
documenting Alban Stewart's botanical specimens collected on the 1905–1906
California Academy of Sciences Galápagos Expedition. One card per specimen with
species, locality on the islands, collection date and field notes; companion to
Stewart's expedition journal.
Harvested from the 6 IA items
csfa788562626 +
csfa788562626Alpha +… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-cas-galapagos-stewart-specimens.Induction-Cooker-Ceramic-Panel-Crack-Identification-Dataset
Induction Cooker Ceramic Panel Crack Identification Dataset
In the current industrial field, the crack problem of induction cooker ceramic panels poses a threat to product safety, leading to potential explosion risks. Existing detection methods mostly rely on manual inspection, which is inefficient and prone to errors. This dataset aims to provide high-quality crack image data to train machine learning models, automating the detection process and improving detection efficiency and… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Induction-Cooker-Ceramic-Panel-Crack-Identification-Dataset.index-cards-peabody-newspaper
Peabody Newspaper Index Cards (Peabody Institute Library, MA)
3,694 typewritten index cards from the Peabody Institute Library — Sutton
Room Local History Resource Center (Peabody, Massachusetts), indexing people,
events, and news in South Danvers / Peabody as recorded in local newspapers.
The information was typed onto cards over decades by library staff as the local
newspaper-of-record archive's principal finding aid.
Plus a companion "Poor Family" genealogy index from the same… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-peabody-newspaper.index-card-blank-content
Index-card blank / content / divider classifier — dataset
Cropped single archival index cards labelled blank, content, or divider, for
training a tiny CPU pre-filter that skips blank/divider cards before expensive VLM metadata
extraction in card-catalogue digitisation pipelines.
Two collections: Boston Public Library (BPL) FRC shelf-list cards and National Library
of Scotland (NLS) Advocates Library cards. Styles differ, so evaluate per collection.
How it was made… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-blank-content.Net.a.Porter.Product.prices.India
Net-a-Porter web scraped data
About the website
In the Asia Pacific region, particularly India, retail industries are witnessing a significant digital transformation. The Ecommerce industry is particularly flourishing, revolutionised by advanced technology, the proliferation of smartphones, and improved internet infrastructure. The online fashion retail sector is a key player in this surge, making a strong foothold in the world of Ecommerce. Net-a-Porter, a premier luxury… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Net.a.Porter.Product.prices.India.IIIT-INDIC-HW-WORDS-Hindi
IIIT-INDIC-HW-WORDS-Hindi
Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images.
Overview
The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research… See the full description on the dataset page: https://huggingface.co/datasets/HarishBonu/IIIT-INDIC-HW-WORDS-Hindi.industrial-defect-dataset
Synthetic Industrial Material Defect Dataset (10k)
Dataset Summary
This dataset contains 5,000 highly detailed, synthetically generated images of various industrial materials exhibiting different types of surface defects. It is designed to be used for training machine learning models in computer vision, specifically for quality control, manufacturing defect detection, and surface anomaly recognition.
All images were generated using Stable Diffusion XL (SDXL) to… See the full description on the dataset page: https://huggingface.co/datasets/himanshu1257/industrial-defect-dataset.index-cards-navy-nurse-corps
US Navy Nurse Corps — Index Cards
25 biographical service cards from the US Navy's Bureau of Medicine and Surgery
(BUMED) History Office, one card per nurse, harvested from their Internet Archive
collection. Includes members of the "Sacred Twenty" — the first female Navy
nurses, appointed in 1908 (Esther V. Hasson, Lena S. Higbee, Elizabeth Hewitt,
Della Knight, M. Estelle Hine, Mary Du Bose, Margaret Murray, Sara Cox, Sara Myer,
Ada M. Pendleton, Florence T. Milburn, Clare L. De… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-navy-nurse-corps.index-cards-southborough-vital-records
Southborough Town Clerk — Vital Records & Veteran Index Cards (MA)
8,661 cards from the Southborough (Massachusetts) Town Clerk office, covering:
Death Index Cards, 1850–2015 — the town clerk's running death index covering ~165 years of Southborough deaths, one card per decedent with surname/given-name/date.
Veteran Card Index + Veteran Grave Registration Card Index — companion indices to Southborough's veteran-affairs records, indexing veterans buried in town cemeteries.… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-southborough-vital-records.enhanced-indian-food-classification
Enhanced Indian Food Classification Dataset
A comprehensive dataset for Indian food classification with 15,404 images across 43 classes.
Dataset Structure
dataset/
├── train/ # Training images
├── validation/ # Validation images
└── test/ # Test images
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("SohlHealth/enhanced-indian-food-classification")
# Access splits
train_data = dataset['train']
val_data… See the full description on the dataset page: https://huggingface.co/datasets/Prateek-044/enhanced-indian-food-classification.Inductor-and-Transformer-Detection-Dataset
Inductor and Transformer Detection Dataset
The current industrial sector faces significant challenges in the accurate inspection of power modules and magnetic components, which directly impacts product quality and reliability. Existing solutions often rely on manual inspection methods that are time-consuming and prone to human error. This dataset aims to address the technical issue of automated defect detection in inductors and transformers by providing a rich set of annotated… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Inductor-and-Transformer-Detection-Dataset.microscopy-datasets-index
Microscopy Datasets Index
I kept losing track of which microscopy datasets exist and what format they're in. So I made an index. 300+ open datasets, searchable by domain, task, microscopy type, and license.
What's in here
A single reference file (JSONL and Parquet) cataloging 309 publicly available microscopy datasets. Each entry includes:
id — short slug
name — human-readable dataset name
source — who published it (Broad Institute, Kaggle, ISBI, etc.)
url — direct link… See the full description on the dataset page: https://huggingface.co/datasets/Laborator/microscopy-datasets-index.
