datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GSM-Symbolic
GSM-Symbolic
This project accompanies the research paper, GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.
Getting Started
In our first release, we provide data for GSM-Symbolic, GSM-Symbolic-P1, and GSM-Symbolic-P2 variants. For each variant, we have released both templates (available on Github), and a sample of generated data that can be used for evaluation.
To load the data, you can use the following code. Note that in… See the full description on the dataset page: https://huggingface.co/datasets/apple/GSM-Symbolic.declref-01-declref_01_symbols-2B
declref-01-declref_01_symbols-2B
Procedurally generated decl-ref-01 documents — the scoped declare/reference language of declref, with noisy references, sparse part-transition chains, interleaved openings, and periodic topic shifts tuned so that a model trained on it matches natural-language / code entropy dynamics (positional entropy profile, its fluctuation texture, and the entropy-quantile distribution), and holds that match as training doubles. Each document is a random… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/declref-01-declref_01_symbols-2B.symbolic-instruction-tuning
Symbolic Instruction Tuning
This is the offical repo to host the datasets used in the paper From Zero to Hero: Examining the Power of Symbolic Tasks in Instruction Tuning. The training code can be found in here.
ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
This dataset is associated with the paper "ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark".
Abstract
Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization with genuine reasoning. To address this gap, we present ASyMOB, a high-resolution dataset of 35,368 validated symbolic math problems spanning… See the full description on the dataset page: https://huggingface.co/datasets/Shalyt/ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark.nanoim-symbolic
nanoIM Symbolic Temporal Aliasing Dataset
nanoIM is a small, synthetic, symbolic dataset for studying temporal aliasing in interaction models. Paired examples can share the same flattened transcript while requiring different target actions because timing, overlap, visual cues, policy events, or background/tool results differ.
Files
File
Hub config
Purpose
data/mini/{train,validation,test}.jsonl
mini
Quick smoke suite for training and evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/jlov7/nanoim-symbolic.Symbolic_Collection
Symbol-LLM: Towards Foundational Symbol-centric Interface for Large Language Models
Paper Link: https://arxiv.org/abs/2311.09278
Project Page: https://xufangzhi.github.io/symbol-llm-page/
🔥 News
🔥🔥🔥 We have made a part of the Symbolic Collection public, including ~88K samples for training (10% of the whole collection). The whole collection is expected to release upon acceptance of the paper.
🔥🔥🔥 The model weights (7B / 13B) are released !
Note
This… See the full description on the dataset page: https://huggingface.co/datasets/Symbol-LLM/Symbolic_Collection.DD_sans_symbolegsm8k-symbolic
GSM8k Symbolic
This is an uploaded form of the dataset from Diffusion of Thoughts,
adapted to follow the same format as Tulu datasets.
Citation
If you find this work useful, please cite the original work:
@article{ye2024diffusion,
title={Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models},
author={Ye, Jiacheng and Gong, Shansan and Chen, Liheng and Zheng, Lin and Gao, Jiahui and Shi, Han and Wu, Chuan and Li, Zhenguo and Bi, Wei and Kong… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/gsm8k-symbolic.DD_avec_symboleDD_mixte_symbolearayun_173-system-law-symbolic-causal-coherence
[DOI] https://doi.org/10.5281/zenodo.17186989
ARAYUN_173 – A System Law for Symbolic and Causal Coherence
Corresponding author: ARAYUN_173 (Independent Research)
E-mail: arayun173 [at] proton [dot] me
Website: arayun173.com
Date: September 2025
Audit Marker: SHA-256(ARAYUN_173|2025-09-04|Draft1)
Contact: arayun173 [at] proton [dot] me
ARAYUN_173 – A System Law for Symbolic and Causal Coherence
Abstract
ARAYUN_173 is not a concept but a system law. It establishes… See the full description on the dataset page: https://huggingface.co/datasets/ARAYUN173/arayun_173-system-law-symbolic-causal-coherence.gene-symbols-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
gene_symbols
This dataset consists of a collection of human gene symbols, including well-known entries like VEGFA, TNF, and BRCA1. Each sample represents a single gene identifier formatted as a standard uppercase text string. The data appears to be a curated list of significant genes often associated with cancer research or cellular signaling pathways.
Dataset size
There are 70 data… See the full description on the dataset page: https://huggingface.co/datasets/joduor/gene-symbols-v1.stateful-gsm-symbolicsymbolic-data-assets-sa
symbolic-data assets (CC BY-SA 4.0 catalogs)
Share-alike-licensed benchmark catalogs for the symbolic-data package, kept SEPARATE from the
permissive catalogs in psaegert/symbolic-data-assets. Everything in this repo is licensed
CC BY-SA 4.0 (see LICENSE: full legal code + derivation statement).
Derived from The On-Line Encyclopedia of Integer Sequences (OEIS, (c) The OEIS Foundation Inc.,
CC-BY-SA-4.0; per-entry A-number attribution in catalog meta) and Wikipedia ("List of… See the full description on the dataset page: https://huggingface.co/datasets/psaegert/symbolic-data-assets-sa.gene-symbols
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
gene_symbols
This dataset consists of a collection of human gene symbols, including well-known entries like VEGFA, TNF, and BRCA1. Each sample represents a single gene identifier formatted as a standard uppercase text string. The data appears to be a curated list of significant genes often associated with cancer research or cellular signaling pathways.
Dataset size
There are 70 data… See the full description on the dataset page: https://huggingface.co/datasets/joduor/gene-symbols.loremipsum-29k-symbolsGSM-Symbolic-TTT
GSM-Symbolic
Dataset Description
This dataset contains symbolic variations of grade-school math word problems.The dataset is constructed by merging multiple generated datasets where each instance corresponds to a symbolic template used to produce variations of a math reasoning problem.
Each instance contains a math word problem along with its corresponding solution and final numeric answer.
Dataset Structure
Data Instances
Each row in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/GSM-Symbolic-TTT.symbolic_sentiment_v1mistral_symbolicLogic_5_7_9_shortchanged_symbolsIOAI_NLP_Training_01_symbolic_nli_private
Symbolic NLI Private Answers
Instructor-only answer key. Upload this directory only to a private Hugging
Face dataset repository.
IOAI_NLP_Training_01_symbolic_nli_public
Symbolic NLI Exercise Inputs
Public evaluation prompts for the Symbolic NLI notebook. Labels are stored in
the separate private instructor payload.
GSM-Symbolic-Synthetic-Dataset-Strict-RegexSymbolicCode-CoTnl2sql-1229-10knous-symbolic-discovery-100k
nous-symbolic-discovery-100k
This dataset was created using the Claude Dataset Skill.
nazy-symbols-classification-openclip-encoded-image-dataGSM-Symbolic-Synthetic-Datasetsymbolic_sentiment_v1_alpacanl2sql_new
