datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TabularMath
📊 TabularMath
TabularMath is a tabular mathematical reasoning benchmark introduced in TabularMath: Understanding Math Reasoning over Tables with Large Language Models. It is built via AUTOT2T, a neuro-symbolic pipeline that automatically transforms math word problems into verified tabular reasoning tasks, enabling scalable evaluation without manual table annotation.
TabularMath jointly assesses reasoning accuracy, information retrieval over complex table structures, and… See the full description on the dataset page: https://huggingface.co/datasets/kevin715/TabularMath.SVAMP_de
SVAMP_de: A Translated German Math Dataset
This is a high-quality German translation of the SVAMP dataset.
We employed a State-of-the-Art (SOTA) LLM within a strict validation pipeline to ensure 100% numerical consistency and logical fidelity.
Math word problems rely on precise numbers and logic. Standard translations often hallucinate numbers or mangle units.
SVAMP_de was generated with a "Strict Logic" pipeline:
Translation: Using SOTA LLMs.
Validation: Every single row was… See the full description on the dataset page: https://huggingface.co/datasets/tabularisai/SVAMP_de.qwen-tabular-data-blindspots
Blind Spots of Qwen/Qwen3.5-4B on Tabular Dataset
Overview
This tabular dataset contains the model inputs, model outputs, expected results for the systematic failure cases of the base model Qwen/Qwen3.5-4B.
The model has 4-billion parameters pretrained language model released within the last six months and fall between the 0.6-6B model requirements for the project.
The model were developed to performs well on general text tasks, it was not trained for structured or… See the full description on the dataset page: https://huggingface.co/datasets/olalytics/qwen-tabular-data-blindspots.
