datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
This dataset is associated with the paper "ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark".
Abstract
Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization with genuine reasoning. To address this gap, we present ASyMOB, a high-resolution dataset of 35,368 validated symbolic math problems spanning… See the full description on the dataset page: https://huggingface.co/datasets/Shalyt/ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark.MGSM-Symbolic
MGSM-Symbolic
MGSM-Symbolic is a multilingual symbolic variant of the Multilingual Grade School Math Benchmark (MGSM).It contains mathematically structured word problems across multiple languages, paired with numerical solutions.
The dataset is designed to support research in:
Multilingual reasoning
Cross-lingual generalisation
Symbolic numerical problem solving
Evaluation of reasoning consistency across languages
Each language contains the same set of problems translated and… See the full description on the dataset page: https://huggingface.co/datasets/lrana/MGSM-Symbolic.nanochat-depo-l0-symbolic-20260715
Nanochat Depo-L0: symbolic
This is a diagnostic, separately versioned Depo source. Each row contains one
16-node cycle and eight queries at depths 1, 2, 4, and 8. Only the eight
single-letter answers and terminal token are supervised. It is designed for a
one-document-per-sequence training protocol and must not be treated as public
Depo v3 data.
GSM-Symbolic-TTT
GSM-Symbolic
Dataset Description
This dataset contains symbolic variations of grade-school math word problems.The dataset is constructed by merging multiple generated datasets where each instance corresponds to a symbolic template used to produce variations of a math reasoning problem.
Each instance contains a math word problem along with its corresponding solution and final numeric answer.
Dataset Structure
Data Instances
Each row in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/GSM-Symbolic-TTT.qwen-blindspot-symbolic-reasoning
Qwen3.5-0.8B Blind Spot Dataset
Multi-Step Symbolic Reasoning Under Linguistic Camouflage
Overview
This dataset documents a targeted blind spot of Qwen/Qwen3.5-0.8B, a compact 0.8-billion parameter instruction-tuned language model released by the Qwen team (Alibaba Cloud) in March 2026.
Final score: 5 / 10 correct.
The central finding is not that the model cannot reason — it visibly tries on every single probe, and gets the arithmetic right more often than… See the full description on the dataset page: https://huggingface.co/datasets/ahmad0999/qwen-blindspot-symbolic-reasoning.
