datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems.
The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.ArtiBench
ArtiBench: Artifact Detection Benchmark
Dataset Structure
Artifact-positive samples:
{
"id": "3qotz3zm",
"has_artifacts": true,
"explanation": "The image presents an aerial view of downtown Manhattan with an unusual twist. A large Ferris wheel, reminiscent of the Millennium Wheel, is oddly positioned next to the skyscrapers, appearing to be fused with the buildings below. ...",
"bboxes": [[114, 253, 432, 694]]
}
Artifact-negative samples:
{
"id": "nkzk0lqs"… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/ArtiBench.european_art
Dataset Card for DEArt: Dataset of European Art
Dataset Summary
DEArt is an object detection and pose classification dataset meant to be a reference for paintings between the XIIth and the XVIIIth centuries. It contains more than 15000 images, about 80% non-iconic, aligned with manual annotations for the bounding boxes identifying all instances of 69 classes as well as 12 possible poses for boxes identifying human-like objects. Of these, more than 50 classes are cultural… See the full description on the dataset page: https://huggingface.co/datasets/biglam/european_art.indian-traditional-artificial-jewellery
Traditional and Handmade Indian Jewellery Dataset
This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks.
Dataset Overview
This dataset is designed to provide detailed information about Indian jewelry… See the full description on the dataset page: https://huggingface.co/datasets/Coder-Dragon/indian-traditional-artificial-jewellery.indian-traditional-artificial-jewellery
Traditional and Handmade Indian Jewellery Dataset
This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks.
Dataset Overview
This dataset is designed to provide detailed information about Indian… See the full description on the dataset page: https://huggingface.co/datasets/manidhardevu/indian-traditional-artificial-jewellery.ArtiFact
ArtiFact
ArtiFact is a large-scale multimodal benchmark of museum artwork records with aligned images and structured metadata. It is designed for evaluating metadata extraction, error detection, semantic querying, and multimodal reasoning over cultural-heritage collections.
The dataset combines records from the Rijksmuseum, the Metropolitan Museum of Art (Met), and the Art Institute of Chicago (AIC), with normalized fields for artists, dates, materials, techniques, dimensions… See the full description on the dataset page: https://huggingface.co/datasets/deem-data/ArtiFact.openart-items-artifacts
OpenArt — Items & Artifacts
openart-items-artifacts is the items artifacts subject collection of the OpenArt family
of open, public-domain art datasets: 25,750 works (11,317 paintings/illustrations · 14,216
photographed objects · 217 unclassified), each paired with a structured VLM caption plus
medium, attribution and inscription metadata.
Human-made objects and the decorative arts — vessels, tools, arms and armor, textiles, furniture and ornament — both as physical artifacts… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-items-artifacts.artigo
ARTigo: Social Image Tagging
60,633 digital reproductions of artworks with crowdsourced tags, from ARTigo — a citizen-science project run since 2010 by the Institute of Art History and the Institute of Informatics at LMU Munich. Players are shown an image and type tags against a clock, scoring when a tag matches one their anonymous opponent gives or one recorded in an earlier session. The aggregate of those game rounds is this dataset. Built from the v1.5 Zenodo deposit (1… See the full description on the dataset page: https://huggingface.co/datasets/biglam/artigo.openbrush-religious-art
OpenBrush Religious Art
Religious paintings from OpenBrush-75K — saints, biblical scenes, devotional works.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,119 you actually want.
Why this subset
A coherent visual genre: religious narrative painting from medieval through early modern. Heavy on Renaissance and Baroque eras. Common… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-religious-art.womens-fashion-catalog
Livostyle Women's Fashion Catalog — Open Data
Open, machine-readable, weekly-updated catalog of 2,766+ curated women's fashion
products from Livostyle.com — a US DTC retailer
(Arcada LLC, Delaware). Free under MIT license for AI/LLM training,
recommender systems, fashion NLP research, and multimodal learning.
TL;DR
from datasets import load_dataset
ds = load_dataset("arturayupov/womens-fashion-catalog")
# ds["products"] → 2,766 products
# ds["images"] → 12,978… See the full description on the dataset page: https://huggingface.co/datasets/arturayupov/womens-fashion-catalog.glami-1m-mteb
GLAMI-1M MTEB multimodal classification
This is an MTEB-ready derivative of the official
glami/glami-1m
release for multilingual image+text fashion classification. The source is
pinned at revision befda45d8d4e8b8082bb8a1912d1f9eb9483991c and remains
licensed under Apache-2.0.
Each example contains the official product image, name and description
joined as text, and the official category ID as label. The complete
116,004-row human-labeled test split is unchanged.
To keep… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-mteb.vintage-artworks-60k-captionedThis is a dataset consisting of 60k vintage artworks from the 20th century, consisting of vintage pulp, sci-fi and pinup artworks from that era.
The dataset has short and long captions for each image, as well as resolution information. The large captions (large_caption column) were made with florence-2-large-ft, and then shortened with llama 3 8b (see short_caption column).
aquaveritas-water-stress
AquaVeritas Water Stress Dataset
Classification labels for 1,656 Sentinel-2 satellite observations across 20 global freshwater and saline sites, used to fine-tune LFM2.5-VL-450M for on-board satellite freshwater monitoring.
Built for the Liquid AI x DPhi Space Hackathon: AI in Space (Hack #05).
Dataset Summary
Each observation covers one of 20 monitored water bodies and includes structured classification labels for two zones:
Core zone (15km x 15km centred on the… See the full description on the dataset page: https://huggingface.co/datasets/Arty1001/aquaveritas-water-stress.the_artist_flux_datasetDataset related to
The Artist | Flux edition.
Auto-tagged on Civitai.
the_artist_natlang_dataset
The Artist for Chroma Dataset
It's the same dataset found at
http://huggingface.co/Greyhaws
What's the difference here?
In order to train the LoRA for Chroma, we ran the images under classification once again, this time uses natural language.
IMPORTANT
you might want to verify if it includes "the_artist" tag, and whether remove it , or just keep it, it won't change much.
Disclaimer
This dataset is offered as is.
All images were generated with an ai image model.
