CoolFace
Datasetpublic

katebor/TableEval

TableEval dataset TableEval is developed to benchmark and compare the performance of (M)LLMs on tables from scientific vs. non-scientific sources, represented as images vs. text. It comprises six data subsets derived from the test sets of existing benchmarks for question answering (QA) and table-to-text (T2T) tasks, containing a total of 3017 tables and 11312 instances. The scienfific subset includes tables from pre-prints and peer-reviewed scholarly publications, while the… See the full description on the dataset page: https://huggingface.co/datasets/katebor/TableEval.

sourceHugging Facemitupdated 1y agoView on Hugging Face
6likes205downloads
Dataset Card

TableEval dataset

![GitHub](https://github.com/esborisova/TableEval-Study) ![ACL](https://aclanthology.org/2025.trl-1.10/) ![arXiv](https://arxiv.org/abs/2507.00152)

TableEval is developed to benchmark and compare the performance of (M)LLMs on tables from scientific vs. non-scientific sources, represented as images vs. text. It comprises six data subsets derived from the test sets of existing benchmarks for question answering (QA) and table-to-text (T2T) tasks, containing a total of 3017 tables and 11312 instances. The scienfific subset includes tables from pre-prints and peer-reviewed scholarly publications, while the non-scientific subset involves tables from Wikipedia and financial reports. Each table is available as a PNG image and in four textual formats: HTML, XML, LaTeX, and Dictionary (Dict). All task annotations are taken from the source datasets. Please, refer to the Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data paper for more details.

Overview and statistics

DatasetTaskSourceImageDictLaTeXHTMLXML
ComTQA (PubTables-1M) <img src='https://img.shields.io/badge/arXiv-2024-red'> <a href='https://arxiv.org/abs/2406.01326'><img src='https://img.shields.io/badge/PDF-blue'></a> <a href='https://huggingface.co/datasets/ByteDance/ComTQA'><img src='https://img.shields.io/badge/Dataset-gold'>VQAPubMed Central⬇️⚙️⚙️⚙️📄
numericNLG <img src='https://img.shields.io/badge/ACL-2021-red'> <a href='https://aclanthology.org/2021.acl-long.115.pdf'><img src='https://img.shields.io/badge/PDF-blue'></a> <a href='https://huggingface.co/datasets/kasnerz/numericnlg?row=0'><img src='https://img.shields.io/badge/Dataset-gold'></a>T2TACL Anthology📄⬇️⚙️⬇️⚙️
SciGen <img src='https://img.shields.io/badge/arXiv-2021-red'> <a href='https://arxiv.org/abs/2104.08296'><img src='https://img.shields.io/badge/PDF-blue'></a> <a href='https://github.com/UKPLab/SciGen/tree/main'><img src='https://img.shields.io/badge/Dataset-gold'></a>T2TarXiv and ACL Anthology📄⬇️📄⚙️⚙️
ComTQA (FinTabNet) <img src='https://img.shields.io/badge/arXiv-2024-red'> <a href='https://arxiv.org/abs/2406.01326'><img src='https://img.shields.io/badge/PDF-blue'></a> <a href='https://huggingface.co/datasets/ByteDance/ComTQA'><img src='https://img.shields.io/badge/Dataset-gold'>VQAEarnings reports of S&P 500 companies📄⚙️⚙️⚙️⚙️
LogicNLG <img src='https://img.shields.io/badge/ACL-2020-red'> <a href='https://aclanthology.org/2020.acl-main.708/'><img src='https://img.shields.io/badge/PDF-blue'></a> <a href='https://huggingface.co/datasets/kasnerz/logicnlg'><img src='https://img.shields.io/badge/Dataset-gold'></a>T2TWikipedia⚙️⬇️⚙️📄⚙️
Logic2Text <img src='https://img.shields.io/badge/ACL-2020-red'> <a href='https://aclanthology.org/2020.findings-emnlp.190/'><img src='https://img.shields.io/badge/PDF-blue'></a> <a href='https://huggingface.co/datasets/kasnerz/logic2text'><img src='https://img.shields.io/badge/Dataset-gold'></a>T2TWikipedia⚙️⬇️⚙️📄⚙️

**Symbol ⬇️ indicates formats already available in the given corpus, while 📄 and ⚙️ denote formats extracted from the table source files (e. g., article PDF, Wikipedia page) and generated from other formats in this study, respectively.

Number of tables per format and data subset
DatasetImageDictLaTeXHTMLXML
ComTQA (PubTables-1M)932932932932932
numericNLG135135135135135
SciGen10351035928985961
ComTQA (FinTabNet)659659659659659
LogicNLG184184184184184
Logic2Text7272727272
Total30173017291029672943
Total number of instances per format and data subset
DatasetImageDictLaTeXHTMLXML
ComTQA (PubTables-1M)62326232623262326232
numericNLG135135135135135
SciGen10351035928985961
ComTQA (FinTabNet)28382838283828382838
LogicNLG917917917917917
Logic2Text155155155155155
Total1131211312112051126211238

Structure

├── ComTQA │ ├── FinTabNet │ │ ├── comtqafintabnet.json │ │ ├── comtqafintabnetimgs.zip │ ├── PubTab1M │ │ ├── comtqapubtab1m.json │ │ ├── comtqapubtab1mimgs.zip │ ├── Logic2Text │ │ ├── logic2text.json │ │ ├── logic2textimgs.zip │ ├── LogicNLG │ │ ├── logicnlg.json │ │ ├── logicnlgimgs.zip │ ├── SciGen │ │ ├── scigen.json │ │ ├── scigenimgs.zip │ ├── numericNLG │ │ ├── numericnlg.json └── └── └── numericnlgimgs.zip

For more details on each subset, please, refer to the respective README.md files: ComTQA, Logic2Text, LogicNLG, SciGen, numericNLG.

Citation

@inproceedings{borisova-etal-2025-table,
    title = "Table Understanding and (Multimodal) {LLM}s: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data",
    author = {Borisova, Ekaterina  and
      Barth, Fabio  and
      Feldhus, Nils  and
      Abu Ahmad, Raia  and
      Ostendorff, Malte  and
      Ortiz Suarez, Pedro  and
      Rehm, Georg  and
      M{\"o}ller, Sebastian},
    editor = "Chang, Shuaichen  and
      Hulsebos, Madelon  and
      Liu, Qian  and
      Chen, Wenhu  and
      Sun, Huan",
    booktitle = "Proceedings of the 4th Table Representation Learning Workshop",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.trl-1.10/",
    pages = "109--142",
    ISBN = "979-8-89176-268-8",
    abstract = "Tables are among the most widely used tools for representing structured data in research, business, medicine, and education. Although LLMs demonstrate strong performance in downstream tasks, their efficiency in processing tabular data remains underexplored. In this paper, we investigate the effectiveness of both text-based and multimodal LLMs on table understanding tasks through a cross-domain and cross-modality evaluation. Specifically, we compare their performance on tables from scientific vs. non-scientific contexts and examine their robustness on tables represented as images vs. text. Additionally, we conduct an interpretability analysis to measure context usage and input relevance. We also introduce the TableEval benchmark, comprising 3017 tables from scholarly publications, Wikipedia, and financial reports, where each table is provided in five different formats: Image, Dictionary, HTML, XML, and LaTeX. Our findings indicate that while LLMs maintain robustness across table modalities, they face significant challenges when processing scientific tables."
}

Funding

This work has received funding through the DFG project NFDI4DS (no. 460234259).

<div style="position: relative; width: 100%;"> <img src="NFDI4DS.png" alt="drawing" width="200" style="position: absolute; bottom: 0; right: 0;"/> </div>