datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dwd-icon-global
Dataset Card for DWD ICON Global Forecast
This dataset is comprised of forecasts from the German Weather Service's (DWD) ICON-Global model from March 2023 to the present with all variables included. Each forecast runs up to 4 days into the future, and the model is ran 4 times per day. This data is an archive of
the publicly available data at https://opendata.dwd.de/weather/nwp/, converted to Zarr format with Xarray. No other processing of the data is performed.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/openclimatefix/dwd-icon-global.dwd-icon-eu
Dataset Card for DWD ICON-EU Forecast
[!NOTE]This HF dataset is no longer being updated.
It has been superseded by the live-updating Zarr of ICON-EU available on Dynamical.org!
And, DWD ICON-EU GRIB2 files are now being archived on Source Co-Op by Dynamical.org and Open Climate Fix.
This dataset is comprised of forecasts from the German Weather Service's (DWD) ICON-EU model.
From 2020-01-01 to March 2023 this archive contains a subset of the ICON-EU variables. From March 2023 to… See the full description on the dataset page: https://huggingface.co/datasets/openclimatefix/dwd-icon-eu.IconStack-48M-Rendered-Trainsto-icon-datasetIconArt🖼️ The dataset IconArt dataset was introduced in the following paper : "Weakly Supervised Object Detection in Artworks" Gonthier et al. ECCV 2018 Workshop Computer Vision for Art Analysis - VISART 2018.
This datasest is designed to evaluate Weakly Supervised object detection methods in paintings.
You can also find project page for the paper here.
This dataset contains 5955 images (from WikiCommons) : a train set of 2978 images and a test set of 2977 images (for classification task). 1480 of… See the full description on the dataset page: https://huggingface.co/datasets/NGonthier/IconArt.win-tiles-icons
Dataset Details
Dataset Description
The pictures were taken from the Discord server https://discord.gg/VMz3GD4d
Relevance as of 29.11.2024
I tried to classify some parts, but I have clumsy and crooked paws to make a proper classifier for all this.
ICON-QA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of ICONQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2021iconqa,
title = {IconQA: A New Benchmark for Abstract Diagram Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ICON-QA.IconQAretro-icon-rl
RetroIcon-RL: 32x32 Retro Pixel-Art Icon Generation with Deterministic Verification
This repository implements RetroIcon-RL — training and evaluating small code models with Reinforcement Learning (RL) and programmatic verification to generate consistent, crisp retro pixel-art icon packs from natural language prompts using sharp SVG block geometry.
🎯 Phase 1 — The 32x32 Task Specification
Canvas: Exactly $32 \times 32$ integer grid (viewBox="0 0 32 32").… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/retro-icon-rl.icon_arm-rolloutsIconAssetsios-app-icons
IOS App Icons
Overview
This dataset contains images and captions of iOS app icons obtained from the iOS Icon Gallery. Each image is paired with a generated caption using a Blip Image Captioning model. The dataset is suitable for image captioning tasks and can be used to train and evaluate models for generating captions for iOS app icons.
Images
The images are stored in the 'images' directory, and each image is uniquely identified with a filename (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/ppierzc/ios-app-icons.svg-icons
Dataset Card for svg-icons
Dataset Description
This dataset contains SVG code examples for training and evaluating SVG models for image vectorization.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
Filename
Unique ID for each SVG
Svg
SVG code
Usage
from datasets import load_dataset
dataset = load_dataset("starvector/svg-icons")… See the full description on the dataset page: https://huggingface.co/datasets/starvector/svg-icons.icongenai-svg-captions
IconGenAI SVG Captions
Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models.
Part of the IconGenAI research project.
Files
Two files are provided at different stages of the processing pipeline:
File
Records
Purpose
icons_captioned_merged.jsonl
275,912
Full license-filtered corpus with VLM-generated captions and collection metadata
icons_training_captioned.jsonl227,821
Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.MMSVG-IconOmniSVG: A Unified Scalable Vector Graphics Generation Model
![Project Page]
Dataset Card for MMSVG-Icon
Dataset Description
This dataset contains SVG icon examples for training and evaluating SVG models for text-to-SVG and image-to-SVG task.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
id
Unique ID for each SVG
svg
SVG code (resized to 200×200, simplified with picosvg)
description… See the full description on the dataset page: https://huggingface.co/datasets/OmniSVG/MMSVG-Icon.3d_icon
3D icons Dataset
This dataset contains free-licensed images, downloaded from unsplash. Curated and created by:
Maria Shalabaieva
Alexander Shatov
pid-icons-mergedsvg-icons-simple
Dataset Card for svg-icons-simple
Dataset Description
This dataset contains SVG code examples for training and evaluating SVG models for image vectorization.
Dataset Structure
Features
The dataset contains the following fields:
Field Name
Description
Filename
Unique ID for each SVG
Svg
SVG code
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/starvector/svg-icons-simple.easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter
easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter
Augmented version of datasets/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP with coordinate jitter.
For each original example, 1 additional copies were created. Each copy
randomly jitters the target coordinate by ±1 pixel in both X and Y. The
assistant coordinate in messages is updated, and bbox/normalized_bbox
are shifted when present.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter.easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP
easyr1-63k-hard-qwen7b-easy-gta1-nores-jedi-fix-synced-ui-vision-grounding-pro-apps-manually-labeled-icon-data-from-yt-4MP
Merged dataset composed of the following sources:
/Users/anasawadalla/Desktop/easyr1-57k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced (57011 samples in split train)
ui-vision-grounding-4MP (5790 samples in split train)
easyr1-v2-pro-apps-manually-labeled-icon-data-from-yt-4MP (230 samples in split train)
Summary
Generated on: 2025-09-10… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP.IconQA_litebrill_iconclass
Dataset Card for Brill Iconclass AI Test Set
Dataset Summary
A test dataset and challenge to apply machine learning to collections described with the Iconclass classification system.
This dataset contains 87749 images with Iconclass metadata assigned to the images. The iconclass metadata classification system is intended to provide 'the comprehensive classification system for the content of images.'.
Iconclass was developed in the Netherlands as a standard… See the full description on the dataset page: https://huggingface.co/datasets/biglam/brill_iconclass.ICON_Materialmm_iconqaeasyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-answer-keyiconclass-vlmLMA-indivudal-projectpi05-taco-libero32-evals
pi0.5 TACO LIBERO-32 Evaluations
This repository contains the evaluation artifacts for the 36 checkpoints in
Berkeley-ICON-Lab/pi05-taco-libero32-checkpoints.
The final protocol uses five random seeds (7, 17, 27, 37, 47) and ten fixed
LIBERO initial states per task and seed. The repository includes:
Final CSV, TSV, and JSON summaries with seed-level sample standard deviation
and 95% confidence intervals.
The experiment manifest and checkpoint mapping.
Per-task and per-seed raw… See the full description on the dataset page: https://huggingface.co/datasets/Berkeley-ICON-Lab/pi05-taco-libero32-evals.iconclass-vlm-sfticonclass-vlm-brillfull
Iconclass VLM — brill full labels
Training-ready VLM iconclass-classification dataset rebuilt from the fuller, cleaner
source labels in biglam/brill_iconclass
(CC0). Recovers labels lost to truncation in davanstrien/iconclass-vlm-sft.
Source images: same Brill Arkyves images as biglam/brill_iconclass, bytes passed through verbatim (no re-encode).
Labels: full Iconclass codes with operators (+n), key-combos :, and qualifiers (TEXT) kept intact. Empty/sentinel tokens stripped; ~5… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/iconclass-vlm-brillfull.
