datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TabMWPSelectionThis dataset is a high-fidelity selection from the Tabular Math Word Problems (TabMWP) benchmark (Lu et al., 2023). TabMWP is a leading resource for evaluating mathematical reasoning over heterogeneous tabular and textual data. To address potential noise and ensure the highest standards of logical grounding, this curated version consists of 100 hand-verified examples. Each entry has been audited to confirm that the multi-step reasoning chains—including information look-up and numerical… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/TabMWPSelection.uae-sales-table-qa
🇦🇪 UAE Sales Table QA (Arabic)
⚠️ Note: All data in this dataset is synthetically generated using random values for learning and experimentation purposes. It does not represent real-world business data.
🧠 Overview
UAE Sales Table QA (Arabic) is an Arabic Question–Answering dataset for table reasoning and data analysis, generated from 21 UAE-style CSV tables.Each example includes:
Question — a natural-language query about the data
Steps — human-readable reasoning… See the full description on the dataset page: https://huggingface.co/datasets/zSynctic/uae-sales-table-qa.snowfox-financial-table-data
snowfox-financial-table-data
SnowFox — financial-table extraction training (snowfox_financial_table tasks).
Contents
train.jsonl (1680 rows)
validation.jsonl (162 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for SnowFox (Michael Anthony Falabella).
chartqa-tables
ChartQA Tables
This dataset contains pre-extracted tables and metadata from the ChartQA dataset by Ahmed Masry et al.
Dataset Description
ChartQA is a benchmark for question answering about charts with visual and logical reasoning. This companion dataset provides:
Structured tables extracted from chart images (CSV format)
Formatted tables in the paper's format for model input
Purpose
The original ChartQA paper evaluated models in two modes:
With gold tables… See the full description on the dataset page: https://huggingface.co/datasets/nmayorga7/chartqa-tables.guru-table-verl
Guru Table VERL
This dataset contains 8,230 table reasoning samples from 3 datasets (HiTab, MultiHierTT, FinQA) for reinforcement learning training with VERL (Volcano Engine Reinforcement Learning). The data is extracted and preprocessed from LLM360/guru-RL-92k.
Dataset Summary
Guru is a reasoning model trained using cross-domain reinforcement learning. This dataset focuses on table reasoning tasks where models must analyze hierarchical tables and financial data to answer… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/guru-table-verl.markdown-table-expert
Markdown Table Expert
A large-scale dataset for teaching language models to read, understand, and reason over markdown tables. Contains 44,000 samples (40,000 train + 4,000 validation) spanning 35 real-world domains with detailed step-by-step reasoning traces.
Why This Dataset
Markdown tables are everywhere — in documentation, reports, READMEs, financial statements, and web content. Yet most LLMs struggle with structured tabular data, especially when asked to perform… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-expert.table-r1-zero-verl
Table-R1-Zero (VERL Format)
This dataset contains 69,265 table reasoning problems from the Table-R1-Zero-Dataset, converted to VERL (Volcano Engine Reinforcement Learning) format for reinforcement learning training workflows.
Source: Table-R1/Table-R1-Zero-Dataset
License: Apache 2.0
Note: System prompts have been removed from all examples for better compatibility with other VERL datasets. The dataset now contains only user messages with table reasoning problems. Ground truth… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/table-r1-zero-verl.hecm-plf-tables
HECM Principal Limit Factor (PLF) Tables
Complete PLF lookup table with 4,864 age-rate combinations for calculating reverse mortgage borrowing capacity. Ages 62-99, interest rates 3.0%-18.875%.
How PLF Works
PLF determines what percentage of your home's value you can access:
Higher age = higher PLF (more money available)
Lower interest rate = higher PLF
Example: Age 72 at 6.5% rate = PLF of 0.434 (43.4% of home value)
Data Fields
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/hecm-plf-tables.
