datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IconStack-48M-Rendered-TrainIconArt🖼️ The dataset IconArt dataset was introduced in the following paper : "Weakly Supervised Object Detection in Artworks" Gonthier et al. ECCV 2018 Workshop Computer Vision for Art Analysis - VISART 2018.
This datasest is designed to evaluate Weakly Supervised object detection methods in paintings.
You can also find project page for the paper here.
This dataset contains 5955 images (from WikiCommons) : a train set of 2978 images and a test set of 2977 images (for classification task). 1480 of… See the full description on the dataset page: https://huggingface.co/datasets/NGonthier/IconArt.ICON-QA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of ICONQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2021iconqa,
title = {IconQA: A New Benchmark for Abstract Diagram Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ICON-QA.IconQAicon_arm-rolloutssvg-icons
Dataset Card for svg-icons
Dataset Description
This dataset contains SVG code examples for training and evaluating SVG models for image vectorization.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
Filename
Unique ID for each SVG
Svg
SVG code
Usage
from datasets import load_dataset
dataset = load_dataset("starvector/svg-icons")… See the full description on the dataset page: https://huggingface.co/datasets/starvector/svg-icons.ios-app-icons
IOS App Icons
Overview
This dataset contains images and captions of iOS app icons obtained from the iOS Icon Gallery. Each image is paired with a generated caption using a Blip Image Captioning model. The dataset is suitable for image captioning tasks and can be used to train and evaluate models for generating captions for iOS app icons.
Images
The images are stored in the 'images' directory, and each image is uniquely identified with a filename (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/ppierzc/ios-app-icons.icongenai-svg-captions
IconGenAI SVG Captions
Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models.
Part of the IconGenAI research project.
Files
Two files are provided at different stages of the processing pipeline:
File
Records
Purpose
icons_captioned_merged.jsonl
275,912
Full license-filtered corpus with VLM-generated captions and collection metadata
icons_training_captioned.jsonl227,821
Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.MMSVG-IconOmniSVG: A Unified Scalable Vector Graphics Generation Model
![Project Page]
Dataset Card for MMSVG-Icon
Dataset Description
This dataset contains SVG icon examples for training and evaluating SVG models for text-to-SVG and image-to-SVG task.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
id
Unique ID for each SVG
svg
SVG code (resized to 200×200, simplified with picosvg)
description… See the full description on the dataset page: https://huggingface.co/datasets/OmniSVG/MMSVG-Icon.pid-icons-merged3d_icon
3D icons Dataset
This dataset contains free-licensed images, downloaded from unsplash. Curated and created by:
Maria Shalabaieva
Alexander Shatov
svg-icons-simple
Dataset Card for svg-icons-simple
Dataset Description
This dataset contains SVG code examples for training and evaluating SVG models for image vectorization.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
Filename
Unique ID for each SVG
Svg
SVG code
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/starvector/svg-icons-simple.easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter
easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter
Augmented version of datasets/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP with coordinate jitter.
For each original example, 1 additional copies were created. Each copy
randomly jitters the target coordinate by ±1 pixel in both X and Y. The
assistant coordinate in messages is updated, and bbox/normalized_bbox
are shifted when present.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter.IconQA_liteeasyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP
easyr1-63k-hard-qwen7b-easy-gta1-nores-jedi-fix-synced-ui-vision-grounding-pro-apps-manually-labeled-icon-data-from-yt-4MP
Merged dataset composed of the following sources:
/Users/anasawadalla/Desktop/easyr1-57k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced (57011 samples in split train)
ui-vision-grounding-4MP (5790 samples in split train)
easyr1-v2-pro-apps-manually-labeled-icon-data-from-yt-4MP (230 samples in split train)
Summary
Generated on: 2025-09-10… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP.brill_iconclass
Dataset Card for Brill Iconclass AI Test Set
Dataset Summary
A test dataset and challenge to apply machine learning to collections described with the Iconclass classification system.
This dataset contains 87749 images with Iconclass metadata assigned to the images. The iconclass metadata classification system is intended to provide 'the comprehensive classification system for the content of images.'.
Iconclass was developed in the Netherlands as a standard… See the full description on the dataset page: https://huggingface.co/datasets/biglam/brill_iconclass.mm_iconqaeasyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-answer-keyiconclass-vlmiconclass-vlm-sfticonclass-vlm-brillfull
Iconclass VLM — brill full labels
Training-ready VLM iconclass-classification dataset rebuilt from the fuller, cleaner
source labels in biglam/brill_iconclass
(CC0). Recovers labels lost to truncation in davanstrien/iconclass-vlm-sft.
Source images: same Brill Arkyves images as biglam/brill_iconclass, bytes passed through verbatim (no re-encode).
Labels: full Iconclass codes with operators (+n), key-combos :, and qualifiers (TEXT) kept intact. Empty/sentinel tokens stripped; ~5… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/iconclass-vlm-brillfull.IconQA_filteredeasyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-dense-rewardmteb-nl-iconclass-clsIconclass is a hierarchical system for classifying the subjects and content of artworks. Each artwork can be annotated with one or more Iconclass codes, where each code represents a specific concept depicted in the work. The dataset is sourced from the Netherlands Institute for Art History and includes annotations linked to artwork titles.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-iconclass-cls.iconclip-search-benchmark
IconClip Search Benchmark
A BEIR-format information-retrieval benchmark for icon search by
text intent. 120 hand-curated paraphrase queries against a
22 827-icon corpus spanning 11 open-license icon libraries
(Lucide, Phosphor, Tabler, Heroicons, Bootstrap, Carbon, Font Awesome,
Iconoir, Ionicons, Material Symbols, RemixIcon).
The dataset supports two retrieval tasks on the same queries / qrels:
Task
Document side
What it measures
t2t — text→text
Icon name + tags +… See the full description on the dataset page: https://huggingface.co/datasets/Cortiq-Labs/iconclip-search-benchmark.iconclass_with_splitsiconocracy-corpus
ICONOCRACY Corpus
Release snapshot for warholana/iconocracy-corpus built from the local iconocracy-corpus repository.
Current release
Release tag: 2026-08-13-viewer-fix
Generated at: 2026-08-13T02:08:14Z
Corpus items: 335
Canonical records: 335
Coded items: 286
Countries represented: 19
Mean endurecimento score: 0.964
Master-record schema versions: 1.0, 2.0.0
Contract
This dataset is a release artifact, not a live working mirror.
Source-of-truth… See the full description on the dataset page: https://huggingface.co/datasets/warholana/iconocracy-corpus.Anime_Iconsiconocracy-corpus-sync-2026-09-02
ICONOCRACY Corpus
Unified release snapshot (2026-09-02-sync-fix) that merges GitLab main with Hub-only records.
Current release
Release tag: 2026-09-02-sync-fix
Generated at: 2026-09-02T07:12:58Z
Corpus items: 337
Canonical records: 337
Coded items: 288
Countries represented: 26
Status
This dataset replaces the stale warholana/iconocracy-corpus snapshot (335 rows) until that repo accepts a new release commit.
Release notes
Restore… See the full description on the dataset page: https://huggingface.co/datasets/warholana/iconocracy-corpus-sync-2026-09-02.IconStack-48M-Rendered-Dev
