CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01garrethlee /comprehensive-arithmetic-problemstext1M<n<10M0 likes12k downloads4mo agoHugging Face02garrethlee /comprehensive-arithmetic-problems-carriestext1M<n<10M0 likes7.9k downloads2y agoHugging Face03EleutherAI /arithmeticA small battery of 10 tests that involve asking language models a simple arithmetic problem in natural language.text10K<n<100K5 likes4.5k downloads4y agoHugging Face04Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes3.8k downloads1mo agoHugging Face05Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes1.5k downloads1mo agoHugging Face06ec75hash /qwen36-arithmetic-readouts Arithmetic Intermediate Readout Sensitivity on Qwen3.6-27B Arithmetic intermediate detection with a Jacobian lens depends sharply on the prompt token being read and the numeral forms accepted by the scorer. Across 105 order-of-operations items, the recorded hosted-lens responses contain the intermediate at rank 1 on 48 items at the trailing space, versus 4 at the preceding token. On the 25 held-out items with two-digit intermediates, rank-1 detection falls from 13 to 3 when… See the full description on the dataset page: https://huggingface.co/datasets/ec75hash/qwen36-arithmetic-readouts.tabularothern<1K0 likes1k downloads7d agoHugging Face07garrethlee /simple-arithmetic-problemstext100K<n<1M2 likes767 downloads2y agoHugging Face08SagheerLab /Arithmetic-Reasoning SagheerLab/Arithmetic-Reasoning A high-quality synthetic arithmetic and elementary mathematics reasoning dataset for training and evaluating small language models - not an "ultimate math" claim, but a clean, verified, tiered reasoning dataset where every answer is programmatically checked. This dataset was built to train 100M-ish models that benefit disproportionately from clean, unambiguous examples. At 50M examples (45M train / 2.5M val / 2.5M test, ~5GB parquet) it is… See the full description on the dataset page: https://huggingface.co/datasets/SagheerLab/Arithmetic-Reasoning.tabulartext-generation10M<n<100M3 likes640 downloads27d agoHugging Face09Yujivus /nanochat-climbmix-arithmetic-base6 nanochat ClimbMix + Arithmetic: base-6 numeral world This is a deterministic base-6 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 6. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.texttext-generation10M<n<100M0 likes579 downloads1mo agoHugging Face10mib-bench /arithmetic_additiontabular10K<n<100K0 likes405 downloads1y agoHugging Face11Skanth007 /arithmetic-logical-pointer-setARTHIMETIC-lOGICAL_POINTER_SET 0 likes289 downloads18d agoHugging Face12mib-bench /arithmetic_subtractiontabular10K<n<100K0 likes269 downloads1y agoHugging Face13liodon-ai /nanochat-calendar-arithmetic-base10 nanochat Base-10 Calendar Arithmetic A deterministic, base-10 arithmetic corpus scoped to three cyclic calendar units: hour-of-day (mod 24), day-of-week (mod 7), and month-of-year (mod 12). Companion to Yujivus/nanochat-climbmix-arithmetic-base10, built the same way but scoped to real modular calendar units instead of free-integer add/sub/mul/div/mod. Every example is a single line — question and answer collapsed into one equation, no exposed reasoning: 23:00 + 18965h = 04:00… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/nanochat-calendar-arithmetic-base10.texttext-generation100K<n<1M0 likes268 downloads25d agoHugging Face14shoumenchougou /RWKV-7-ArithmeticRWKV-7-Arithmetic-0.1B 加减法运算模型的训练和测试数据集。 该模型实现基础加减法运算和加减法方程求解功能,能够处理整数部分为 1-12 位、小数部分为 0-6 位的数值,支持中英文数字、全半角格式以及大小写字符的多种表示形式,可实现基础加减法运算和加减法方程求解功能。 训练数据集说明 以下是我们使用的加减法训练数据类型,共包含 30000587 33000147 条单轮加减法 QA 数据,约 1B(1014434168) token。 数据文件名 数据条数 数据说明 示例 ADD_4M 3997733 1. 使用‘全角’、‘中文数字’、‘大写中文数字’随机替换整个数字2. 运算符附近有 1~2 个随机空格3. 含简单自然语言描述/自然语言噪声 {"text": "User: 249476576 减 796580834 还剩多少?\n\nAssistant: -547104258"} ADD_2M 1999673 1. 使用‘全角’、‘中文数字’、‘大写中文数字’随机替换整个数字2. 运算符附近有 1~2… See the full description on the dataset page: https://huggingface.co/datasets/shoumenchougou/RWKV-7-Arithmetic.text10M<n<100M0 likes236 downloads1y agoHugging Face15ESITime /tram-arithmetic-responsestext10K<n<100K0 likes234 downloads1y agoHugging Face16arithmetic-circuit-overloading /synthetic-dataset-1d-500K-50K-0.1-reverse-padzerotext10M<n<100M0 likes170 downloads7mo agoHugging Face17donoway /deepmind-math-arithmetictext10M<n<100M1 likes167 downloads11mo agoHugging Face18arithmetic-circuit-overloading /synthetic-dataset-v2-3d-5M-500K-0.1-padzerotext10M<n<100M0 likes165 downloads6mo agoHugging Face19flexitok /mod-arithmetic Modular Arithmetic Dataset Synthetic dataset of modular-arithmetic problems of the form a mod b, paired with the result and a hypothesis about the most suitable tokenizer. Tokenizer hypothesis For a mod b where b = 2^k × 5^j (no other prime factors), only the rightmost max(k, j) digits of a determine the answer, because 10^max(k,j) ≡ 0 (mod b). A tokenizer that groups digits right-to-left in chunks of that size exposes the relevant information as a single token. For all… See the full description on the dataset page: https://huggingface.co/datasets/flexitok/mod-arithmetic.tabularquestion-answering1M<n<10M0 likes161 downloads7mo agoHugging Face20lintang /numerical_reasoning_arithmetic Generated dataset for testing numerical reasoningtabular1K<n<10K0 likes157 downloads4y agoHugging Face21arithmetic-circuit-overloading /synthetic-dataset-v2-3d-3M-300K-0.1-reversetext10M<n<100M0 likes147 downloads6mo agoHugging Face22Lots-of-LoRAs /task087_new_operator_addsub_arithmetic Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task087_new_operator_addsub_arithmetic Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task087_new_operator_addsub_arithmetic.texttext-generation1K<n<10K0 likes138 downloads2y agoHugging Face23arithmetic-circuit-overloading /synthetic-dataset-v2-3d-5M-500K-0.1-reverse-padzerotext10M<n<100M0 likes130 downloads6mo agoHugging Face24Lots-of-LoRAs /task085_unnatural_addsub_arithmetic Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task085_unnatural_addsub_arithmetic Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task085_unnatural_addsub_arithmetic.texttext-generation1K<n<10K0 likes128 downloads2y agoHugging Face25arithmetic-circuit-overloading /synthetic-dataset-v2-3d-3M-300K-0.1-padzerotext10M<n<100M0 likes124 downloads6mo agoHugging Face26arithmetic-circuit-overloading /synthetic-dataset-1d-500K-50K-0.2-reversetext10M<n<100M0 likes118 downloads7mo agoHugging Face27braindecode /arithmetic_zyma2019 EEG Dataset This dataset was created using braindecode, a deep learning library for EEG/MEG/ECoG signals. Dataset Information Property Value Recordings 72 Type Windowed (from Epochs object) Channels 19 Sampling frequency 200 Hz Total duration 2:22:06 Windows/samples 1,707 Size 1.46 MB Format zarr Quick Start from braindecode.datasets import BaseConcatDataset # Load from Hugging Face Hub dataset =… See the full description on the dataset page: https://huggingface.co/datasets/braindecode/arithmetic_zyma2019.0 likes118 downloads6mo agoHugging Face28thoughtworks /arithmetic-sorl-data Arithmetic SoRL Data Training and evaluation data for the SoRL Arithmetic Interpretability Study. Small transformers trained on integer addition/subtraction, with SoRL to externalize carry/borrow circuits as explicit abstraction tokens. Reference: Quirke et al., "Understanding Addition and Subtraction in Transformers" (2024). Paper: arXiv:2402.02619 — see Table 8 for complexity classification and Section 3 for sub-task definitions. Dataset Structure Subfolder… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/arithmetic-sorl-data.tabulartext-generation1M<n<10M0 likes116 downloads5mo agoHugging Face29neurallambda /arithmetic_dataset Arithmetic Puzzles Dataset A collection of arithmetic puzzles with heavy use of variable assignment. Current LLMs struggle with variable indirection/multi-hop reasoning, this should be a tough test for them. Inputs are a list of strings representing variable assignments (c=a+b), and the output is the integer answer. Outputs are filtered to be between [-100, 100], and self-reference/looped dependencies are forbidden. Splits are named like: train_N 8k total examples of puzzles with N… See the full description on the dataset page: https://huggingface.co/datasets/neurallambda/arithmetic_dataset.text100K<n<1M0 likes115 downloads2y agoHugging Face30arithmetic-circuit-overloading /synthetic-dataset-2d-1M-100K-0.2-reversetext10M<n<100M0 likes113 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.