psp-dada/ChartArena
ChartArena A Comprehensive Bilingual Benchmark for General Chart Parsing across Families, Scenarios, and Formats Github Repo • Paper Overview ChartArena is a bilingual benchmark for evaluating the chart parsing capabilities of vision-language models. It covers the full difficulty spectrum of real-world charts, spanning eight chart families across three visual scenarios and two languages. Contents Statistics Dataset… See the full description on the dataset page: https://huggingface.co/datasets/psp-dada/ChartArena.
ChartArena <!-- omit in toc -->
A Comprehensive Bilingual Benchmark for General Chart Parsing across Families, Scenarios, and Formats
<p align="center"> <a href="https://github.com/pspdada/ChartArena">Github Repo</a> • <a href="https://arxiv.org/abs/2606.01348">Paper</a> </p>
Overview <!-- omit in toc -->
ChartArena is a bilingual benchmark for evaluating the chart parsing capabilities of vision-language models. It covers the full difficulty spectrum of real-world charts, spanning eight chart families across three visual scenarios and two languages.
<table align="center"> <p align="center"> <img src="docs/figures/ChartArena_overview.jpg" width="80%" /> </p> </table>
Contents <!-- omit in toc -->
Statistics
Chart families
Visual scenarios
Languages
Bilingual: Chinese (ZH) and English (EN)
Dataset Structure
data/
├── ChartArena.jsonl # annotations for all samples
└── images/Data Format
Each line of ChartArena.jsonl is a JSON object:
{
"img_path": "images/xxx.png",
"chart_type": "柱状图",
"img_type": "电子印刷",
"lang_type": "中文",
"anno": "..."
}Usage
Please refer to the Github repository for inference and evaluation scripts.
# Quick start
git clone https://github.com/pspdada/ChartArena
cd ChartArena
pip install -r requirements.txt
# Run inference
python infer.py --api_type openai_compat --model_name <model> --base_url <url>
# Score
python judge.py
# Generate analysis report
python analyze.pyCitation
@article{peng2026chartarena,
title = {{ChartArena}: Benchmarking Chart Parsing across Languages, Scenarios, and Formats},
author = {Peng, Shangpin and Li, Gengluo and Wan, Xingyu and Zhang, Chengquan and Feng, Hao and Wu, Binghong and Shen, Huawen and Wang, Weinong and Cai, Ziyi and Tian, Zhuotao and Hu, Han and Ma, Can and Zhou, Yu},
journal = {arXiv preprint arXiv:2606.01348},
year = {2026}
}License
This dataset is released for research purposes only.
