datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VIZWIZ_TRAIN_data_with_imagesdata-viz-qa
DataVizQA
This is a image-text dataset for question answering over data visualizations.
This dataset is a random sampling of the following datasets:
18% from FigureQA
64% from PlotQA
28% from ChartQA
orbura-dataviz-dataset
Orbura AutoScientist Data Visualization Dataset
Dataset Description
A synthetic instruction-tuning corpus for data visualization tasks, built for
the AutoScientist Challenge
Part 2 (Data Visualization). Every example is deterministically generated —
no human annotation, no LLM-generated labels — so the ground truth is exact
and reproducible.
Task types
Task
Weight
Description
chart_qa
45%
Arithmetic, comparison, and trend questions about… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/orbura-dataviz-dataset.dataviz-html-dataset
DataViz HTML Dashboard Dataset
100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
data/*.csv — CSV data files used by 15 dashboards
README.md — Dataset card
Columns
Column
Type
Description
filename
string
File name of the dashboard
title
string
Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.dataviz-charts-html-dataset
DataViz Individual Charts HTML Dataset
120 individual chart HTML files — 12 chart types × 10 themes — for teaching models to generate themed single-chart visualizations.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
README.md — Dataset card
Columns
Column
Type
Description
filename
string
treemap-tech-innovation.html
title
string
e.g. "Bar Chart — Midnight Galaxy"
chart_type
string
bar, line, area, pie… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-charts-html-dataset.orbura-dataviz-augmentedautoscientist-dataviz-dataset
AutoScientist adapted dataset — dataviz
Adaption Labs AutoScientist v5 adapted fine-tuning data for the dataviz category.
dataviz_adapted.jsonl — prompt/completion pairs used for QLoRA SFT.
dataviz_v5_raw.csv — full Adaption output (prompt, completion, enhanced_prompt, chosen, rejected, reasoning_trace, embeddings) used for DPO.
Paired weights: Rishidar/autoscientist-dataviz-qlora (Kaggle mirror rishidard/autoscientist-dataviz-qlora).
orbura-autoscientist-dataviz-multimodal-pilotadaption-africa-data-viz-instruct
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-africa_data_viz_instruct
This instruction-tuning dataset contains 184 samples focused on data visualization within an African context, covering demographics, infrastructure, and economic indicators. The entries feature prompts and completions that explain statistical trends, recommend chart types, and provide Python code using libraries like Matplotlib and GeoPandas.… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-africa-data-viz-instruct.MedsiML_EN_HIN_Data
MedSiML
Dataset Description
This dataset contains simplified English and simplified Hindi sentence pairs derived from the MedSiML dataset. The data has been filtered using the Cynical Data Selection algorithm to retain examples that are most representative of the target distribution while reducing redundancy.
Source Dataset
This dataset is derived from the MedSiML dataset:
MedSiML: A Multilingual Approach for Simplifying Medical Texts… See the full description on the dataset page: https://huggingface.co/datasets/vishnu-vizz/MedsiML_EN_HIN_Data.dataviz-99-nl2vis-dataset
DataViz-99: NL2VIS Dataset Card
Dataset Description
DataViz-99 is a closed-label instruction-tuning dataset for natural language to visualization (NL2VIS) tasks. Built from real benchmark data (nvBench + Spider), it teaches models to convert natural language queries into structured chart specifications.
Dataset Summary
Metric
Value
Total Examples
17,457
Format
JSONL (prompt/completion pairs)
Languages
English, Hindi, Spanish, French… See the full description on the dataset page: https://huggingface.co/datasets/pandeyankit84/dataviz-99-nl2vis-dataset.dataviz-sample
Dataset Card for "dataviz-sample"
More Information needed
VIZWIZ_VALIDATION_data_with_imagesadaption-dataviz-expert-recommendations
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-dataviz_expert_recommendations
This dataset contains expert-level data visualization recommendations generated for diverse analytical scenarios across multiple domains. Each entry includes a specific prompt detailing the domain, available fields, and constraints, paired with a rigorous response specifying chart types, encodings, and design justifications. The content focuses on best… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-dataviz-expert-recommendations.world_data_vizWMT_EXT_DATAViz_Lang_Synthetic_Data_0Viz_Lang_Real_Data_one_thirdViz_Lang_Real_DataViz_Lang_Synthetic_Data_2Viz_Lang_test_DataViz_Lang_Synthetic_Data_1data_viz_datasetofa-viz-data
