CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceM4 /ChartQA Dataset Card for "ChartQA" More Information needed image10K<n<100K68 likes23k downloads3y agoHugging Face02ChartGalaxy /ChartGalaxy ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation 🤗 Dataset | 🖥️ Code | 📄 Paper | 📄 Arxiv 🔥 News [2026.09] 🎉🎉 A new high-quality batch of 14,809 synthetic infographic charts has been added. This update features more complex layouts and richer chart variations. [2026.02] 🎉🎉 A new batch of data has been added, comprising 108,208 infographic charts. This update features broader diversity in title designs and more polished layouts… See the full description on the dataset page: https://huggingface.co/datasets/ChartGalaxy/ChartGalaxy.imagevisual-question-answering1K<n<10K89 likes23k downloads16d agoHugging Face03lmms-lab-encoder /ChartQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of ChartQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{masry2022chartqa, title={ChartQA: A benchmark for question answering about charts with visual and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ChartQA.image1K<n<10K26 likes17k downloads3y agoHugging Face04ibm-granite /ChartNet ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding 🌐 Homepage | 📖 arXiv 📝 Changelog June 3, 2026 — Release of grounded_qa subset and completed reasoning subset (both subject to Notice Regarding Data Availability) May 15, 2026 — Added link to 30K real-world charts and detailed captions dataset released by our collaborators Abaka AI/2077AI. April 29, 2026 — Release of an additional 2.5 million row subset core_permissive (subject to… See the full description on the dataset page: https://huggingface.co/datasets/ibm-granite/ChartNet.imageimage-to-text1M<n<10M47 likes15k downloads4mo agoHugging Face05ckchaos /ChartDiff ChartDiff: A Large-Scale Benchmark for Comprehending Pairs of Charts Overview ChartDiff is a large-scale benchmark for cross-chart comparative summarization, designed to evaluate whether vision-language models can identify differences and generate coherent comparative descriptions across pairs of charts. Unlike existing chart understanding datasets that emphasize single-chart interpretation, ChartDiff requires models to compare two charts jointly and generate a concise… See the full description on the dataset page: https://huggingface.co/datasets/ckchaos/ChartDiff.imagesummarization1K<n<10K0 likes12k downloads6mo agoHugging Face06CSU-JPG /Chart2CodeFrom Charts to Code: A Hierarchical Benchmark for Multimodal Models Welcome to Chart2Code! If you find this repo useful, please give a star ⭐ for encouragement. Data Overview Chart2Code is a hierarchical benchmark for evaluating multimodal models on chart understanding and chart-to-code generation. The dataset is organized into five Hugging Face configurations: level1_direct level1_customize level1_figure level2 level3 In the current Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/CSU-JPG/Chart2Code.imageimage-text-to-text1K<n<10K3 likes11k downloads5mo agoHugging Face07Colinyyy /ChartM3image1K<n<10K3 likes6.9k downloads1y agoHugging Face08ChrisFan /ChartDQAimagequestion-answering1K<n<10K2 likes4.2k downloads1y agoHugging Face09ahmed-masry /ChartQAIf you wanna use the dataset, you need to download the zip file manually from the "Files and versions" tab. Please note that this dataset can not be directly loaded with the load_dataset function from the datasets library. If you want a version of the dataset that can be loaded with the load_dataset function, you can use this one: https://huggingface.co/datasets/ahmed-masry/chartqa_without_images But it doesn't contain the chart images. Hence, you will still need to use the images stored in… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/ChartQA.image10K<n<100K32 likes3.7k downloads2y agoHugging Face10NgTMDuc /VLLM_ChartQAtext10K<n<100K0 likes3.4k downloads2y agoHugging Face11surgeai /chartography Chartography Chartography Chartography measures professional visual reasoning over the charts that professionals stake decisions on: Sankey diagrams, candlestick charts, contour maps, back-trajectories, Bode plots, and more. What it tests The benchmark contains 100 real-world prompts and chart images spanning 12 professional domains, including Finance & Investing, Healthcare, Manufacturing & Supply Chain, and STEM fields from Chemistry to Geosciences to Electrical Engineering.… See the full description on the dataset page: https://huggingface.co/datasets/surgeai/chartography.imagevisual-question-answeringn<1K0 likes3.3k downloads2mo agoHugging Face12vinod-anbalagan /adaption-charts-p2-gold Adaption Charts P2 — Gold Chart-QA Dataset A verified, quality-first chart question-answering dataset built for the Adaption Labs AutoScientist Challenge (Part 2, Data Visualization track). Two sources: a programmatically generated synthetic core (correct-by-construction) and a hand-authored hardset built from real public dashboards and reports. At a glance 3803 rows total — 3705 synthetic + 98 hardset 7 chart types — bar, line, grouped_bar, stacked_bar, pie… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/adaption-charts-p2-gold.imagevisual-question-answering1K<n<10K0 likes3k downloads1mo agoHugging Face13lytang /ChartMuseum [NeurIPS 2025] ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models Authors: Liyan Tang, Grace Kim, Xinyu Zhao, Thom Lake, Wenxuan Ding, Fangcong Yin, Prasann Singhal, Manya Wadhwa, Zeyu Leo Liu, Zayne Sprague, Ramya Namuduri, Bodun Hu, Juan Diego Rodriguez, Puyuan Peng, Greg Durrett Leaderboard 🥇 | Paper 📃 | Code 💻 Overview ChartMuseum is a chart question answering benchmark designed to evaluate reasoning capabilities of large… See the full description on the dataset page: https://huggingface.co/datasets/lytang/ChartMuseum.imagequestion-answering1K<n<10K7 likes2.7k downloads1y agoHugging Face14vikhyatk /chartqaimage1K<n<10K0 likes2.2k downloads2y agoHugging Face15papylove /bettor-chart-images0 likes2k downloads10m agoHugging Face16ahmed-masry /ChartQAPro ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering 🤗Dataset | 🖥️Code | 📄Paper The abstract of the paper states that: Charts are ubiquitous, as people often use them to analyze data, answer questions, and discover critical insights. However, performing complex analytical tasks with charts requires significant perceptual and cognitive effort. Chart Question Answering (CQA) systems automate this process by enabling models to interpret and reason with… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/ChartQAPro.textvisual-question-answering1K<n<10K21 likes1.9k downloads1y agoHugging Face17SD122025 /ChartGen-200Kimageimage-to-text100K<n<1M9 likes1.9k downloads1y agoHugging Face18opendatalab /ChartVerse-SFT-600KChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page. This dataset contains non-trivial samples filtered by failure rate (r > 0), ensuring that every sample provides meaningful learning signal. Samples that are too easy (r = 0, where the model always answers correctly) are… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-600K.imagevisual-question-answering100K<n<1M11 likes1.9k downloads8mo agoHugging Face19Peppertuna /ChartQAimagequestion-answeringn<1K5 likes1.6k downloads3y agoHugging Face20openbmb /VisRAG-Ret-Test-ChartQA Dataset Description This is a VQA dataset based on Charts from ChartQA dataset from ChartQA. Load the dataset from datasets import load_dataset import csv def load_beir_qrels(qrels_file): qrels = {} with open(qrels_file) as f: tsvreader = csv.DictReader(f, delimiter="\t") for row in tsvreader: qid = row["query-id"] pid = row["corpus-id"] rel = int(row["score"]) if qid in qrels:… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Test-ChartQA.imagen<1K1 likes1.5k downloads2y agoHugging Face21Evanwu50020 /visualization-chartsimagen<1K0 likes1.4k downloads1y agoHugging Face22ahmed-masry /chartqa_without_images Dataset Card for "chartqa_without_images" If you wanna load the dataset, you can run the following code: from datasets import load_dataset data = load_dataset('ahmed-masry/chartqa_without_images') The dataset has the following structure: DatasetDict({ train: Dataset({ features: ['imgname', 'query', 'label', 'type'], num_rows: 28299 }) val: Dataset({ features: ['imgname', 'query', 'label', 'type'], num_rows: 1920 }) test:… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/chartqa_without_images.text10K<n<100K1 likes1.3k downloads3y agoHugging Face23opendatalab /ChartVerse-SFT-1.8MChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page. This dataset contains all verified correct samples without failure rate filtering. Unlike SFT-600K which excludes easy samples (r=0), SFT-1800K includes the complete set of truth-anchored QA pairs for maximum coverage and scale.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-1.8M.imagevisual-question-answering1M<n<10M139 likes1.3k downloads7mo agoHugging Face240xzanuee /ChartVerse-SFT-1800KChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page. This dataset contains all verified correct samples without failure rate filtering. Unlike SFT-600K which excludes easy samples (r=0), SFT-1800K includes the complete set of truth-anchored QA pairs for maximum coverage and scale.… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/ChartVerse-SFT-1800K.imagevisual-question-answering1M<n<10M1 likes1.2k downloads8mo agoHugging Face25starriver030515 /chartverse-allimage1M<n<10M0 likes1.2k downloads8mo agoHugging Face26psp-dada /ChartArena ChartArena A Comprehensive Bilingual Benchmark for General Chart Parsing across Families, Scenarios, and Formats Github Repo • Paper Overview ChartArena is a bilingual benchmark for evaluating the chart parsing capabilities of vision-language models. It covers the full difficulty spectrum of real-world charts, spanning eight chart families across three visual scenarios and two languages. Contents Statistics Dataset… See the full description on the dataset page: https://huggingface.co/datasets/psp-dada/ChartArena.imageimage-to-text1K<n<10K0 likes1k downloads3mo agoHugging Face27alisalam /ChartGazeimage1K<n<10K4 likes993 downloads1y agoHugging Face28DanhVuiVe /ChartQA_small_preprocessedimage1K<n<10K0 likes937 downloads2y agoHugging Face29Zhihan /Chart2NCode Aligned Multi-View Scripts for Universal Chart-to-Code Generation Chart2NCode is a multi-language chart-to-code dataset: every chart image is paired with three aligned plotting scripts — Python (matplotlib), R (ggplot2), and LaTeX (pgfplots) — that render visually equivalent outputs. Released with the paper "Aligned Multi-View Scripts for Universal Chart-to-Code Generation" (ACL 2026). Paper: https://arxiv.org/abs/2604.24559 Code: https://github.com/zhihan72/CharLuMA Models:… See the full description on the dataset page: https://huggingface.co/datasets/Zhihan/Chart2NCode.imageimage-text-to-text100K<n<1M1 likes903 downloads5mo agoHugging Face30manifesta /scientific-chart-qa-17k Scientific Chart QA, 17,070 rows A multimodal chart-interpretation dataset built around one idea: teaching a model when not to answer matters as much as teaching it to answer. One in seven questions here cannot be answered from its figure, and the correct response is cannot be determined. Baseline vision-language models overwhelmingly guess a plausible-looking number instead. That is the behaviour this set targets. The four things worth… See the full description on the dataset page: https://huggingface.co/datasets/manifesta/scientific-chart-qa-17k.imagevisual-question-answering10K<n<100K0 likes691 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.