datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cola
COLA: Compose Objects Localized with Attributes
Self-contained Hugging Face port of the COLA benchmark from the paper
"How to adapt vision-language models to Compose Objects Localized with Attributes?".
📄 Paper: https://arxiv.org/abs/2305.03689
🌐 Project page: https://cs-people.bu.edu/array/research/cola/
💻 Original code & data: https://github.com/ArijitRay1993/COLA
This repository bundles the benchmark annotations as Parquet files and the referenced
images as regular files… See the full description on the dataset page: https://huggingface.co/datasets/array/cola.smart-bin-detect
arudaev/smart-bin-detect
Training data for Smart Bin Recognition – a validator ("is there a bin?")
and an identifier ("which bin?"). The design lives in docs/04-ml-pipeline.md
in the project repo, which is private; the manifests here carry per-image
provenance and are the authoritative record of what this dataset contains.
Every image carries provenance: source, source URL, licence, region,
capture date, annotator where known, label origin (human / machine /
legacy /… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/smart-bin-detect.Icarus-dataset
Icarus
A unified multi-modal curriculum dataset for evolutionary neural architecture search. Every row is one self-contained Task = {meta, support, query}, where support and query are lists of (input_Field, output_Field) pairs. The inner loop trains on support; fitness is scored on query. Support is non-empty for every task. Encoders read the Field descriptor (axes, value_type, n_classes, value_range, mask); mask is True where a value is padding/ignored. meta.class_names, when… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/Icarus-dataset.artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems.
The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.ArtiBench
ArtiBench: Artifact Detection Benchmark
Dataset Structure
Artifact-positive samples:
{
"id": "3qotz3zm",
"has_artifacts": true,
"explanation": "The image presents an aerial view of downtown Manhattan with an unusual twist. A large Ferris wheel, reminiscent of the Millennium Wheel, is oddly positioned next to the skyscrapers, appearing to be fused with the buildings below. ...",
"bboxes": [[114, 253, 432, 694]]
}
Artifact-negative samples:
{
"id": "nkzk0lqs"… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/ArtiBench.chest-xray-14-320
NIH Chest X-ray14 - 320x320 Processed for CheXVision
Project Resources
GitHub repository
Presentation deck
Live demo
Scratch model
DenseNet model
This dataset repackages the raw NIH Chest X-ray14 source dataset from
alkzar90/NIH-Chest-X-ray-dataset
into a data-only Parquet dataset for the CheXVision project.
Dataset Summary
Source format: 12 ZIP archives of original chest X-ray images plus CSV manifests
Output format: data-only Parquet shards under data/… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/chest-xray-14-320.Ar-MUSA
Data Directory Structure
The Ar-MUSA directory contains annotated datasets organized by batches and annotation teams. Each batch is labeled with a number, and the annotation team is indicated by a letter. The structure is as follows:
Ar-MUSA
├── Annotation 1a
│ ├── frames # Contains the extracted frames for each record
│ ├── audios # Contains the corresponding audio files
│ ├── transcripts # Contains the transcripts of the audio files
│ └── annotations.csv #… See the full description on the dataset page: https://huggingface.co/datasets/Skhaled/Ar-MUSA.european_art
Dataset Card for DEArt: Dataset of European Art
Dataset Summary
DEArt is an object detection and pose classification dataset meant to be a reference for paintings between the XIIth and the XVIIIth centuries. It contains more than 15000 images, about 80% non-iconic, aligned with manual annotations for the bounding boxes identifying all instances of 69 classes as well as 12 possible poses for boxes identifying human-like objects. Of these, more than 50 classes are cultural… See the full description on the dataset page: https://huggingface.co/datasets/biglam/european_art.indian-traditional-artificial-jewellery
Traditional and Handmade Indian Jewellery Dataset
This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks.
Dataset Overview
This dataset is designed to provide detailed information about Indian jewelry… See the full description on the dataset page: https://huggingface.co/datasets/Coder-Dragon/indian-traditional-artificial-jewellery.chest-xray-14
NIH Chest X-ray14 — Processed for CheXVision
This dataset wraps the NIH Chest X-ray14 dataset, preprocessed for the CheXVision project.
Labels
Label
Count
Prevalence
Infiltration
19,894
17.7%
Effusion
13,317
11.9%
Atelectasis
11,559
10.3%
Nodule
6,331
5.6%
Mass
5,782
5.2%
Pneumothorax
5,302
4.7%
Consolidation
4,667
4.2%
Pleural_Thickening
3,385
3.0%
Cardiomegaly
2,776
2.5%
Emphysema
2,516
2.2%
Edema
2,303
2.1%
Fibrosis
1,686
1.5%… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/chest-xray-14.indian-traditional-artificial-jewellery
Traditional and Handmade Indian Jewellery Dataset
This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks.
Dataset Overview
This dataset is designed to provide detailed information about Indian… See the full description on the dataset page: https://huggingface.co/datasets/manidhardevu/indian-traditional-artificial-jewellery.ArtiFact
ArtiFact
ArtiFact is a large-scale multimodal benchmark of museum artwork records with aligned images and structured metadata. It is designed for evaluating metadata extraction, error detection, semantic querying, and multimodal reasoning over cultural-heritage collections.
The dataset combines records from the Rijksmuseum, the Metropolitan Museum of Art (Met), and the Art Institute of Chicago (AIC), with normalized fields for artists, dates, materials, techniques, dimensions… See the full description on the dataset page: https://huggingface.co/datasets/deem-data/ArtiFact.white-cabbage-leaf-damage
White Cabbage Leaf Damage Dataset
Description
This dataset contains images of white cabbage leaves with various types of damage. It is designed for researchers and developers working on agricultural computer vision and plant pathology detection.
Dataset Structure
The dataset is organized into folders representing different classes of leaf damage or healthy states.
Usage
You can use this dataset with the datasets library:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/white-cabbage-leaf-damage.openbrush-religious-art
OpenBrush Religious Art
Religious paintings from OpenBrush-75K — saints, biblical scenes, devotional works.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,119 you actually want.
Why this subset
A coherent visual genre: religious narrative painting from medieval through early modern. Heavy on Renaissance and Baroque eras. Common… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-religious-art.ArMeme
ArMeme Dataset
Overview
ArMeme is the first multimodal Arabic memes dataset that includes both text and images, collected from various social media platforms. It serves as the first resource dedicated to Arabic multimodal research. While the dataset has been annotated to identify propaganda in memes, it is versatile and can be utilized for a wide range of other research purposes, including sentiment analysis, hate speech detection, cultural studies, meme generation, and… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ArMeme.womens-fashion-catalog
Livostyle Women's Fashion Catalog — Open Data
Open, machine-readable, weekly-updated catalog of 2,766+ curated women's fashion
products from Livostyle.com — a US DTC retailer
(Arcada LLC, Delaware). Free under MIT license for AI/LLM training,
recommender systems, fashion NLP research, and multimodal learning.
TL;DR
from datasets import load_dataset
ds = load_dataset("arturayupov/womens-fashion-catalog")
# ds["products"] → 2,766 products
# ds["images"] → 12,978… See the full description on the dataset page: https://huggingface.co/datasets/arturayupov/womens-fashion-catalog.geolayers
Geolayers-Data
-->
This dataset card contains usage instructions and metadata for all data-products released with our paper:Using Multiple Input Modalities can Improve Data-Efficiency and O.O.D. Generalization for ML with Satellite Imagery. We release 3 modified versions of 3 benchmark datasets spanning land-cover segmentation, tree-cover regression, and multi-label land-cover classification tasks. These datasets are augmented with auxiliary, geographic inputs. A full list of… See the full description on the dataset page: https://huggingface.co/datasets/arjunrao2000/geolayers.openart-items-artifacts
OpenArt — Items & Artifacts
openart-items-artifacts is the items artifacts subject collection of the OpenArt family
of open, public-domain art datasets: 25,750 works (11,317 paintings/illustrations · 14,216
photographed objects · 217 unclassified), each paired with a structured VLM caption plus
medium, attribution and inscription metadata.
Human-made objects and the decorative arts — vessels, tools, arms and armor, textiles, furniture and ornament — both as physical artifacts… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-items-artifacts.artigo
ARTigo: Social Image Tagging
60,633 digital reproductions of artworks with crowdsourced tags, from ARTigo — a citizen-science project run since 2010 by the Institute of Art History and the Institute of Informatics at LMU Munich. Players are shown an image and type tags against a clock, scoring when a tag matches one their anonymous opponent gives or one recorded in an earlier session. The aggregate of those game rounds is this dataset. Built from the v1.5 Zenodo deposit (1… See the full description on the dataset page: https://huggingface.co/datasets/biglam/artigo.vibe-landing-page-arena
Vibe Landing Page Arena
A large-scale human preference dataset for evaluating AI-generated landing page design quality. 36,000 pairwise judgments from 3,492 annotators comparing landing pages generated by Claude Code, Cursor, Lovable, and Replit across 100 prompts and 4 design dimensions.
Overview
Metric
Value
Total judgments
36,000
Unique annotators
3,492
Prompts
100
Business categories
97
Design tones
82
Tools compared
4 (Claude Code, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/vibe-landing-page-arena.glami-1m-mteb
GLAMI-1M MTEB multimodal classification
This is an MTEB-ready derivative of the official
glami/glami-1m
release for multilingual image+text fashion classification. The source is
pinned at revision befda45d8d4e8b8082bb8a1912d1f9eb9483991c and remains
licensed under Apache-2.0.
Each example contains the official product image, name and description
joined as text, and the official category ID as label. The complete
116,004-row human-labeled test split is unchanged.
To keep… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-mteb.StairvsNonStair_Dataset
24-679 (Fall 2026): Stairs and Non-Stair Images
ArinRoths/StairvsNonStair_Dataset
Photos of stairs and non-stairs scenes, prepared as square RGB images with multiple separately
generated training variants. The goal of this dataset is to classify whether or not stairs are
present in an image.
Source and task
The original dataset contains 32 images that I collected and organized into stairs and
non_stairs folders. There are 16 original stairs images and 16 original… See the full description on the dataset page: https://huggingface.co/datasets/ArinRoths/StairvsNonStair_Dataset.ArGuard-Task1
ArGuard – Track A: Arabic Hateful Memes
This repository hosts the official dataset for Track A of the
ArGuard shared task: multimodal hateful-meme detection in Arabic.
Each instance is an Arabic meme (image + OCR-extracted overlaid text)
manually annotated for hatefulness and fine-grained sub-types.
Content warning. The dataset contains text and imagery that is
offensive, discriminatory, or otherwise harmful by design. Handle
with care.
Track A subtasks
Given a… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ArGuard-Task1.samuel-and-audrey-photography-metadata-archive
Samuel & Audrey Photography Metadata Archive
This dataset contains a structured metadata archive for the Samuel & Audrey Media Network travel photography collection hosted on SmugMug.
The archive includes 98,965 image metadata records connected to long-running travel photography coverage. Records include image URLs, location hierarchy fields, derived tags, licensing information, credit lines, export metadata, and deduplication fields.
This dataset provides metadata and source URLs… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-photography-metadata-archive.Photo_Dataset
24-679 (Fall 2026): Stairs and Non-Stair Images
ArinRoths/Photo_Dataset
Photos of stairs and non-stairs scenes, prepared as square RGB images with multiple separately
generated training variants. The goal of this dataset is to classify whether or not stairs are
present in an image.
Source and task
The original dataset contains 32 images that I collected and organized into stairs and
non_stairs folders. There are 16 original stairs images and 16 original non-stairs… See the full description on the dataset page: https://huggingface.co/datasets/ArinRoths/Photo_Dataset.vintage-artworks-60k-captionedThis is a dataset consisting of 60k vintage artworks from the 20th century, consisting of vintage pulp, sci-fi and pinup artworks from that era.
The dataset has short and long captions for each image, as well as resolution information. The large captions (large_caption column) were made with florence-2-large-ft, and then shortened with llama 3 8b (see short_caption column).
arithmetic-llm
arithmetic-llm
MNIST 图像 → 字节级算术 训练数据集。
用于教学向的字节级算术大模型 (byte-omni-model-zh):给定一个显式数字 (digit 0-9) 与一张 MNIST 图像 (隐式数字),预测两者之和。
输入: digit_byte + MNIST_image_bytes
输出: result_bytes
示例:
digit = "1" + image(数字 "2") → 预测 "3"
Schema
列
类型
说明
id
int
序号
image_base64
str
MNIST 图像 PNG (base64)
pixels
float[]
28×28 展平像素 (0-1), 784 维
label
int
图像数字标签 (0-9)
split
str
train (59134) / test (9943)
image_hash
str
SHA-256 逐字节指纹
phash
str
16×16 感知哈希位串 (与 label… See the full description on the dataset page: https://huggingface.co/datasets/wcpsoft/arithmetic-llm.brain_tumors_dataset
Brain Tumour Dataset
This dataset contains 4 types of brain's condition gioma, healthy, meningioma and pituitary with around more than 5k images
You can use this dataset for Brain Tumor Classification
Burberry.Product.prices.United.Arab.Emirates
Burberry web scraped data
About the website
The luxury fashion industry in the EMEA region, particularly in the United Arab Emirates, is a robust and rapidly growing market. This growth is primarily fuelled by the affluent consumer base and tourists, who have a keen interest in high-end and prestigious fashion labels. Middle Eastern consumers often look at luxury goods as a status symbol, thus driving up demand for premium brands like Burberry. In recent years, Ecommerce… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Burberry.Product.prices.United.Arab.Emirates.aquaveritas-water-stress
AquaVeritas Water Stress Dataset
Classification labels for 1,656 Sentinel-2 satellite observations across 20 global freshwater and saline sites, used to fine-tune LFM2.5-VL-450M for on-board satellite freshwater monitoring.
Built for the Liquid AI x DPhi Space Hackathon: AI in Space (Hack #05).
Dataset Summary
Each observation covers one of 20 monitored water bodies and includes structured classification labels for two zones:
Core zone (15km x 15km centred on the… See the full description on the dataset page: https://huggingface.co/datasets/Arty1001/aquaveritas-water-stress.
