datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VaaniVAANI is an India-representative multi-modal multi-lingual dataset.
The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages.
From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts.
Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.ids-project-artifactsledger-long-context-multi-kpi
the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks.
OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking.
Dataset Description
This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks.
Configs
Config
Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.artist-styles
artist-styles
Static gallery of artist styles. Plain HTML/JS (index.html) reading from
data/artists.json and images/ — no build step, no dependencies.
Running with Docker
docker.sh runs the site in a python:3.12-slim container serving the project
directory with python3 scripts/server.py (a stdlib-only server: static files
plus a small favorites API). The container is named artist-styles, restarts
automatically (--restart=always), and serves on port 7803 by… See the full description on the dataset page: https://huggingface.co/datasets/jtreminio/artist-styles.relaion-art
Relaion Art - LLM-Annotated
Original Source
📌 Introduction
This dataset comprises images and annotations from the original Relaion Art Dataset.
Out of the 8M images, a subset of 3.66M images has been annotated with automatic methods (Image-text-to-text models).
Captions
The annotations include four annotation columns:
dense_caption: A dense annotation about the image
vqa: Visual Question-Answers related to the image. JSON dictionary embedded as a… See the full description on the dataset page: https://huggingface.co/datasets/Fhrozen/relaion-art.laion-art-en-colorcanny
Dataset Card for "laion-art-en-colorcanny"
More Information needed
ARTO-Gen-Dataset
ARTO-KG: A Synthetic Artwork Dataset for Knowledge-Enhanced Understanding
Dataset Description
ARTO-KG is a large-scale synthetic artwork dataset that bridges visual content and structured knowledge through ontology-guided automated generation. Each artwork is annotated with comprehensive RDF knowledge graphs aligned with the ARTO ontology.
Dataset Summary
Total Artworks: 10,108 high-resolution images (1024×1024)
Object Instances: 39,878 (average… See the full description on the dataset page: https://huggingface.co/datasets/youngcan1/ARTO-Gen-Dataset.sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.PrismLayersPro
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
We introduce PrismLayersPro, a 20K high-quality multi-layer transparent image dataset with rewritten style captions and human filtering.
PrismLayersPro is curated from our 200K dataset, PrismLayers, generated via MultiLayerFLUX.
Dataset Structure
📑 Dataset Splits (by Style)
The PrismLayersPro dataset is divided into 21 splits based on visual style… See the full description on the dataset page: https://huggingface.co/datasets/artplus/PrismLayersPro.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
art-my-vibe-imagesannotated-3DGS-artifacts
Puzzle Similarity
Project page | Paper | Code
This repository contains the dataset presented in the ICCV 2025 paper "Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions"Authors: Nicolai Hermann, Jorge Condor, and Piotr Didyk
Dataset Description
The Dataset consists of 36 hand-selected 3D Gaussian Splatting renderings containing common reconstruction artefacts, (aligned) ground truths, human-annotated… See the full description on the dataset page: https://huggingface.co/datasets/nihermann/annotated-3DGS-artifacts.mvl-sib-sent2img-mteb
MVL-SIB sentence-to-image for MTEB
Native MTEB retrieval packaging of the official MVL-SIB single-reference
(k=1) sentence-to-image task. Each of 205 language subsets has 3,012
sentence queries, four candidate images per query, and one correct image.
All subsets reference one shared 70-image corpus file.
Source and changes
Derived from the official WueNLP/MVL-SIB
dataset and MVL-SIB paper,
pinned at 1df5974e8fb204e91ee70cef2b3b7196a14b390f. The official builder's… See the full description on the dataset page: https://huggingface.co/datasets/artist/mvl-sib-sent2img-mteb.splash-art-gacha-collection-10k
Splash Art Collection 10K
This collection features 11,755 character splash arts or 角色立绘 sourced from 47 gacha games, meticulously gathered from Fandom and Biligame WIKI.
The dataset is suitable for fine-tuning T2I models on splash art generation domain, utilizing the image and prompt fields. It includes a mix of both high- and low-quality splash arts of various styles, allowing you to curate and select the images that best suit your training needs.
Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/splash-art-gacha-collection-10k.indian-art-styles
🎨 Indian Art Styles Dataset
A comprehensive image classification dataset covering 34 distinct Indian painting and art styles with 27,139 images in total. This dataset is designed for training Vision Transformer (ViT) and CNN-based classifiers to recognize traditional Indian art styles.
Dataset Overview
Style
Region
Medium
Image Count
aipan
Uttarakhand
floor/wall painting
6
bengal_school
West Bengal
painting
1287
bhil
Madhya Pradesh / Rajasthan /… See the full description on the dataset page: https://huggingface.co/datasets/Divya0001/indian-art-styles.artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems.
The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
ArtiBench
ArtiBench: Artifact Detection Benchmark
Dataset Structure
Artifact-positive samples:
{
"id": "3qotz3zm",
"has_artifacts": true,
"explanation": "The image presents an aerial view of downtown Manhattan with an unusual twist. A large Ferris wheel, reminiscent of the Millennium Wheel, is oddly positioned next to the skyscrapers, appearing to be fused with the buildings below. ...",
"bboxes": [[114, 253, 432, 694]]
}
Artifact-negative samples:
{
"id": "nkzk0lqs"… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/ArtiBench.ArtemisMix-v1
ArtemisMix-v1
Stage-2 (multimodal instruction fine-tuning) corpus for Artemis, the
Schneewolf Labs vision-language flagship that grafts a Qwen3-VL ViT + MLP
projector onto the A2 decoder (the A3 lineage).
This is the Lite slice (v1): layers L1 + L4 only, 350,000 rows.
The planned L2 (multimodal tool/agent) and L3 (custom distill) layers are
produced separately and concatenated later.
Composition
Layer
Rows
Share
Purpose
L1 general multimodal instruction… See the full description on the dataset page: https://huggingface.co/datasets/schneewolflabs/ArtemisMix-v1.real-fake-ai-generated-art-images
🎨 Real and Fake (AI-Generated) Art Images Dataset
21,642 balanced images — 10,821 real artworks and 10,821 AI-generated
images — for training models to distinguish authentic art from GAN-generated fakes.
🧭 Overview
This dataset is part of the FauxFinder project, designed to build
advanced models capable of distinguishing between authentic artworks
and AI-generated images. Ideal for binary classification, GAN research,
and computer vision benchmarking.… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/real-fake-ai-generated-art-images.anime-art-curated
anime-art-curated
A curated WebDataset of anime / digital illustration images with Danbooru-style
text tags, intended as a clean training corpus for text-to-image and
illustration-style model fine-tuning.
Provenance
This dataset is a filtered re-publication of an earlier dataset
(advokat/artist390k, now deleted) which was itself algorithmically curated
from broader booru sources by selecting popular images. The original source
inadvertently contained material that the… See the full description on the dataset page: https://huggingface.co/datasets/advokat/anime-art-curated.rgb_articubotdataset for rgb articubot
sdxl_images_easy_prompts-artists-seed1contract_parse
data/ — new-run outputs
Parser/verification runs in THIS project write here (e.g. data/runs/...).
The frozen regression baseline — the past parsed files the drift-check in
step (c) compares against — lives in the sibling repo and is reached via the
baseline/ symlink: baseline/data/runs/turnNN_*/.
szl-artifacts
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZL Artifacts — Build Artifact Registry
Artifact boundary - audited 2026-07-15: this Hugging Face dataset repository
is a mixed 106-file, 36,100,909-byte publication/build mirror at revision
91bcb443857f2884ef2bfabaaa6bfdc606c7134a. It is not a uniform table of
DSSE envelopes, and… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-artifacts.ArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/Thunderbolt215215/ArtiMuse-10K.deepseek-ocr-artifacts-test-XXPrismLayersPlus
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
PrismLayersPro100k
PrismLayersPro100k is a high-quality dataset designed for training and evaluating multi-layer transparent image generation models. It contains 100,000 samples, each with multiple layers including foreground objects and background scenes. The dataset aims to facilitate research in transparent object compositing, layer decomposition, and text-aware image generation.… See the full description on the dataset page: https://huggingface.co/datasets/artplus/PrismLayersPlus.hf-assets
