datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svgrepo
Dataset Card for SVGRepo Icons
Dataset Summary
This dataset contains a large collection of Scalable Vector Graphics (SVG) icons sourced from SVGRepo.com. The icons cover a wide range of categories and styles, suitable for user interfaces, web development, presentations, and potentially for training vector graphics or icon classification models. Each icon is provided under a specific open-source or permissive license, clearly indicated in its metadata. The SVG… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/svgrepo.svgfind
Dataset Card for SVGFind Icons
Dataset Summary
This dataset contains a large collection of Scalable Vector Graphics (SVG) icons sourced from SVGFind.com. The icons cover a wide range of categories and styles, suitable for user interfaces, web development, presentations, and potentially for training vector graphics or icon classification models. Each icon is provided under either a Creative Commons license or is in the Public Domain, as clearly indicated in its… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/svgfind.svgfind
Dataset Card for SVGFind Icons
Dataset Summary
This dataset contains a large collection of Scalable Vector Graphics (SVG) icons sourced from SVGFind.com. The icons cover a wide range of categories and styles, suitable for user interfaces, web development, presentations, and potentially for training vector graphics or icon classification models. Each icon is provided under either a Creative Commons license or is in the Public Domain, as clearly indicated in its… See the full description on the dataset page: https://huggingface.co/datasets/wapiuk/svgfind.icongenai-svg-captions
IconGenAI SVG Captions
Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models.
Part of the IconGenAI research project.
Files
Two files are provided at different stages of the processing pipeline:
File
Records
Purpose
icons_captioned_merged.jsonl
275,912
Full license-filtered corpus with VLM-generated captions and collection metadata
icons_training_captioned.jsonl227,821
Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.clker-svg
Dataset Card for Clker.com SVG Images
Dataset Summary
This dataset contains 255,758 public domain SVG vector clipart images collected from Clker.com. Clker.com hosts user-shared vector clip art that is explicitly released into the public domain (CC0). The dataset includes the SVG content itself along with metadata such as titles and tags associated with each image. The SVG files in this dataset have been minified using tdewolff/minify to reduce file size while… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/clker-svg.pelican-svg-drawings
Pelican SVG: frontier model drawings, scored
139 SVGs produced by seven frontier models answering Simon Willison's prompt,
"generate an SVG of a pelican riding a bicycle", each with the score the
pelican_svg_env
OpenEnv environment gave it. Four of the seven are open weights and three are closed.
Simon has run that prompt against nearly every model release since early 2025, but the
results live as embedded images across 129 blog posts and scattered gists. This dataset
exists… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/pelican-svg-drawings.SVG-Sophia
Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning
🧩 SVG-Sophia Dataset
The SVG-Sophia dataset is available at Hugging Face.
After downloading and extraction, the files are organized as follows:
File
Description
cot_img2svg_sft.jsonlCoT training data for the SFT stage — Image-to-SVG task
cot_text2svg_sft.jsonl
CoT training data for the SFT stage — Text-to-SVG task
cot_refinement_sft.jsonl
CoT training data… See the full description on the dataset page: https://huggingface.co/datasets/InternSVG/SVG-Sophia.pokemon-llm-svg-bench-results
Pokémon LLM SVG Bench (V1.1)
This is a snapshot of scored SVG drawings from the Pokémon LLM SVG Bench — LLMs try to draw Pokémon in SVG, and we score the results. Prefer reading the live site for context; for scoring rules and methodology, see About.
Unofficial fan / research project. Pokémon images are copyrighted by The Pokémon Company and related companies involved in developing, operating, and managing the Pokémon series. Pokémon descriptions are sourced from… See the full description on the dataset page: https://huggingface.co/datasets/haxfenx/pokemon-llm-svg-bench-results.svgen-500k-instruct
SVGen Vector Images Dataset Instruct Version
Overview
SVGen is a comprehensive dataset containing 300,000 SVG vector codes from a diverse set of sources including SVG-Repo, Noto Emoji, and InstructSVG. The dataset aims to provide a wide range of SVG files suitable for various applications including web development, design, and machine learning research.
Data Fields
{
"text": "<s>[INST] Icon of Look Up here are the inputs Look Up [/INST] \\n <?xml… See the full description on the dataset page: https://huggingface.co/datasets/umuthopeyildirim/svgen-500k-instruct.text-to-svg
Text-to-SVG Dataset
Overview
This dataset is curated to support training and evaluating large language models (LLMs) for text-to-SVG generation tasks.It combines multiple high-quality sources to provide a diverse and comprehensive collection of SVG code examples paired with textual prompts and structured instructions.The focus is on enabling models to generate standards-compliant SVG graphics from descriptive language.
Dataset Composition
1️⃣ Visual… See the full description on the dataset page: https://huggingface.co/datasets/wexhi/text-to-svg.diverse-svg-prompts
Diverse SVG Prompts
Diverse SVG Prompts is a public collection of 20,000 high-quality,
generated and filtered English briefs for SVG and vector-graphics generation.
It contains 18,000 general illustration prompts and 2,000 lettering prompts.
Schema
The dataset intentionally has only two columns:
prompt: the complete visual brief.
type_tags: a list of category, author-model, and processing tags.
Example:
{
"prompt": "A moonlit mechanical heron..."… See the full description on the dataset page: https://huggingface.co/datasets/Nbardy/diverse-svg-prompts.svg-multimodal-rubrics
SVG Multimodal Rubrics
A multimodal dataset of SVG code generation samples with natural language descriptions and evaluation rubrics. Each sample pairs a detailed prompt (Markdown) with its corresponding SVG source code, covering animations, 3D scenes, games, and visual effects.
Designed for training and evaluating models on visual code generation — generating complex, interactive SVG artwork from natural language descriptions.
Overview
Item
Details
Samples… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/svg-multimodal-rubrics.yupp-svg-20251204
Yupp SVG Dataset: Exploration of the Reasoning and Coding Abilities of Frontier Models
1. Overview
This dataset contains organic user preferences collected from Yupp, where users compare side-by-side AI-generated SVG outputs.
Each entry represents a user chat with model responses and preference comparisons, providing a direct evaluation of model reasoning and coding capabilities through the lens of SVG generation.
This release represents only a small fraction of the SVG… See the full description on the dataset page: https://huggingface.co/datasets/yupp-ai/yupp-svg-20251204.svg-instructshaped-svgs-autocaptioned-1675So, whilst working on my SVG model, I saw this dataset and thought, despite it's relatively small size, it would be much better and much more helpful if the images were captioned. Now, this is V0.1. V. 0.1. As in, it's not done yet. The images have (mostly) i'd say approximately 70-80% acurately captioned. As soon as my server is refilled, i'm going through it with a VQA and an actual vision model, i've a few in mind. Also special prompts to ensure EVERY image is accurate.
In the meantime… See the full description on the dataset page: https://huggingface.co/datasets/MrOvkill/shaped-svgs-autocaptioned-1675.UAV-SVGsvgSVG_diagramssvg-chart-training-data
svg-chart-training-data
915 training examples for SVG business chart generation, across 12 chart types.
Generated using gemma3:12b via Ollama, validated through a 3-gate pipeline (extract → XML parse → coordinate bounds), and used to fine-tune per-type LoRA adapters on Gemma 3 12B via MLX.
Schema
Each line is a JSON object with two fields:
{
"input": {
"chart_type": "bar",
"heading": "Quarterly Revenue ($M)",
"value_format": "dollar",
"data_points":… See the full description on the dataset page: https://huggingface.co/datasets/John-Williams-ATL/svg-chart-training-data.svg_pathDataset of googlefont charset convert to svg
Dataset made with font forge
goodsmash-nft-art-svgsoilguardsvg-stack-qwen-relabel-sharegpt
SVG Stack Qwen Relabel ShareGPT
ShareGPT-format conversation dataset with train, validation, and test splits.
Each record contains a conversations field.
valteu-svg-dataset_subset
Valteu SVG Dataset (Subset)
Text-to-SVG dataset subset.
Splits
train.jsonl
test.jsonl
Fields
gt_svg: ground-truth SVG string
caption_idefics3_short: text description
SVG-CompBench
