datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChartNet
ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding
🌐 Homepage | 📖 arXiv
📝 Changelog
June 3, 2026 — Release of grounded_qa subset and completed reasoning subset (both subject to Notice Regarding Data Availability)
May 15, 2026 — Added link to 30K real-world charts and detailed captions dataset released by our collaborators Abaka AI/2077AI.
April 29, 2026 — Release of an additional 2.5 million row subset core_permissive (subject to… See the full description on the dataset page: https://huggingface.co/datasets/ibm-granite/ChartNet.ChartAnno
ChartAnno: Benchmarking Multimodal Large Language Models for Chart Annotation Generation
The official dataset repository of ChartAnno
1,200 real-world charts · 3,600 instructions · 10,800 instances (+720 D3/SVG) · 3 representations · 17 chart types
1. Data Overview
Annotations are essential to communicative visualization, helping explain data, emphasize key findings… See the full description on the dataset page: https://huggingface.co/datasets/chartanno/ChartAnno.ChartMimic
ChartMimic: Evaluating LMM’s Cross-Modal Reasoning Capability via Chart-to-Code Generation
This is the official dataset repository of ChartMimic.
Kind Note: ChartMimic has been integrated into VLMEvalKit. Welcome to use ChartMimic through VLMEvalKit! Special thanks to the VLMEvalKit team.
1. Data Overview
ChartMimic aims at assessing the visually-grounded code generation capabilities of large multimodal models (LMMs). ChartMimic utilizes information-intensive visual… See the full description on the dataset page: https://huggingface.co/datasets/ChartMimic/ChartMimic.ChartInt
ChartInt
ChartInt is a multimodal chart dataset for chart reconstruction, chart editing, style transfer, interaction editing, and data-update tasks. The Hugging Face release is packaged as a datasets-compatible Parquet dataset so that the Dataset Viewer can display rows, text/code fields, and chart screenshots directly.
Dataset Structure
The dataset contains 2,905 rows in the train split.
Task
Rows
Description
native_reconstruction
556
Reconstruct chart code… See the full description on the dataset page: https://huggingface.co/datasets/xilinghuiye/ChartInt.chartforge
ChartForge: 117,600 table-to-chart instructions, every code row executed
A programmatically generated instruction set for teaching small models to turn a data
table into working chart code, into an exact answer, or into a declarative chart
spec. Built for the Adaption AutoScientist Challenge, Part 2 (Data Visualization);
the request phrasings were co-optimized with Adaptive Data (Adaption Labs).
What this dataset proves, and how you check it
rows
117… See the full description on the dataset page: https://huggingface.co/datasets/Jainamshahhh/chartforge.ChartSense_8645_web_enriched
ChartSense 8645 web enriched
This is a supervised fine tuning dataset that teaches a small language model to
behave like a data analyst on chart and data visualization work, i.e.
it should be able to see through the user's words,
resolve underspecified asks,
push back if required,
critique flawed charts, and
answer the question the chart is a means to.
Built for the AutoScientist Challenge by Adaption Labs.
This is the web-enriched build. The web slice is grown to its full… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense_8645_web_enriched.CharToM-QA
Dataset Card for CharToM-QA
Dataset Details
CharToM-QA is a benchmark introduced in the paper The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters. It comprises 1,035 Theory of Mind (ToM) questions based on characters from classic novels. The benchmark is designed to evaluate ToM-related question-answering (QA) capabilities about characters in the context of novels. In CharToM-QA, the task takes the form of ToM… See the full description on the dataset page: https://huggingface.co/datasets/ZeroXeno/CharToM-QA.ChartSense-10k
ChartSense
This is a supervised fine tuning dataset that teaches a small language model to behave like a data analyst on chart and data visualization work, i.e.
it should be able to see through the user's words,
resolve underspecified asks,
push back if required,
critique flawed charts, and
answer the question the chart is a means to.
Built for the AutoScientist Challenge by Adaption Labs.
What it is
10,038 conversations between a person with data and an… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense-10k.ChartSense_8827_corpus_enriched
ChartSense 8827 corpus enriched
This is a supervised fine tuning dataset that teaches a small language model to
behave like a data analyst on chart and data visualization work, i.e.
it should be able to see through the user's words,
resolve underspecified asks,
push back if required,
critique flawed charts, and
answer the question the chart is a means to.
Built for the AutoScientist Challenge by Adaption Labs.
This is the corpus-enriched build. The corpus slice is grown to its… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense_8827_corpus_enriched.chart-reasoning-mix-v1
Chart Reasoning Mix v1
Training data for fine-tuning compact LLMs (Phi-3 Mini, Qwen 2.5 3B)
to map (natural-language question + SQL result schema) to a
storytelling-grade chart specification.
Part of the SQL Agent LLMOps project.
Total
Sources
Storytelling fields
35,167 rows
2 (nvBench real + OpenAI synth)
chart_type, encoding, title, sort, color_strategy, rationale
Part of the SQL Agent LLMOps project
Dataset
ModelRole… See the full description on the dataset page: https://huggingface.co/datasets/DanielRegaladoCardoso/chart-reasoning-mix-v1.dv-misleading-charts
ChartLieDetector
ChartLieDetector is an original, provenance-tracked instruction-tuning dataset for visualization practitioners, data journalists, analysts, and model builders who need to audit chart integrity without a multimodal pipeline. It pairs serialized chart specifications and captions with construction-grounded distortion labels, minimal path-level corrections, claim support judgments, and plain-language descriptions of the data. No upstream dataset contains these exact… See the full description on the dataset page: https://huggingface.co/datasets/0xkamal7/dv-misleading-charts.svg-chart-render-v1
SVG Chart Render Mix v1
Training data for fine-tuning a small code model (DeepSeek Coder 1.3B)
to map (chart specification JSON) to inline SVG code.
Part of the SQL Agent LLMOps project.
Total
Sources
Input
Output
~25,000 rows
2
structured JSON chart spec
rendered SVG string
Part of the SQL Agent LLMOps project
Dataset
Model
Role
DanielRegaladoCardoso/text-to-sql-mix-v2
Qwen 2.5 Coder 7B
NL question to SQL… See the full description on the dataset page: https://huggingface.co/datasets/DanielRegaladoCardoso/svg-chart-render-v1.ChartSense_11k
ChartSense 11k
This is a supervised fine tuning dataset that teaches a small language model to
behave like a data analyst on chart and data visualization work, i.e.
it should be able to see through the user's words,
resolve underspecified asks,
push back if required,
critique flawed charts, and
answer the question the chart is a means to.
Built for the AutoScientist Challenge by Adaption Labs.
This is the full build: both slices grown to their current size. The two partial… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense_11k.ChartSense
ChartSense
This is a supervised fine tuning dataset that teaches a small language model to behave like a data analyst on chart and data visualization work, i.e.
it should be able to see through the user's words,
resolve underspecified asks,
push back if required,
critique flawed charts, and
answer the question the chart is a means to.
Built for the AutoScientist Challenge by Adaption Labs.
What it is
6,038 conversations between a person with data and an… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/ChartSense.dataviz-charts-html-dataset
DataViz Individual Charts HTML Dataset
120 individual chart HTML files — 12 chart types × 10 themes — for teaching models to generate themed single-chart visualizations.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
README.md — Dataset card
Columns
Column
Type
Description
filename
string
treemap-tech-innovation.html
title
string
e.g. "Bar Chart — Midnight Galaxy"
chart_type
string
bar, line, area, pie… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-charts-html-dataset.text-to-chart-spec-dataset
Text-to-Chart-Spec Dataset
96% Win Rate on Adaption AutoScientist Challenge
Overview
This dataset was created for the Adaption AutoScientist Challenge in the Data Visualization and Chart Interpretation category.
Dataset Statistics
Metric
Value
Total Examples
18,003
Format
JSONL (prompt/completion)
Task
Text to Chart Specification
Win Rate
96%
Output Format
The model outputs chart specifications in a structured… See the full description on the dataset page: https://huggingface.co/datasets/pandeyankit84/text-to-chart-spec-dataset.qwen-qwen2-vl-7b-instruct__vision-understanding-chart-qa-mini__019e3b8585be
Qwen/Qwen2-VL-7B-Instruct on vision.understanding.chart-qa-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
5
N Ok
5
Ok Rate
1
Accuracy
0.6
Accuracy P05
0
Accuracy P50
1
Accuracy P95
1
TTFT P50
84.8863
ms
Total P50 Ms
84.8863
Tokens Out Total
13
Run configuration
Model: Qwen/Qwen2-VL-7B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-vl-7b-instruct__vision-understanding-chart-qa-mini__019e3b8585be.svg-chart-training-data
svg-chart-training-data
915 training examples for SVG business chart generation, across 12 chart types.
Generated using gemma3:12b via Ollama, validated through a 3-gate pipeline (extract → XML parse → coordinate bounds), and used to fine-tune per-type LoRA adapters on Gemma 3 12B via MLX.
Schema
Each line is a JSON object with two fields:
{
"input": {
"chart_type": "bar",
"heading": "Quarterly Revenue ($M)",
"value_format": "dollar",
"data_points":… See the full description on the dataset page: https://huggingface.co/datasets/John-Williams-ATL/svg-chart-training-data.ChartEdit_cot
STEM Image Chain-of-Thought Edit Analysis Dataset
This dataset contains AI-generated Chain-of-Thought (CoT) reasoning for STEM image editing tasks, providing step-by-step analysis of edit operations.
Dataset Structure
The dataset is organized in batches:
Total batches: 2
Each batch is stored in a separate directory (batch_0000, batch_0001, etc.)
Fields
Each item contains:
Source image and caption (from previous stage)
Edit command (from original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/ChartEdit_cot.
