datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
countdown-backtrackingStep Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
train: 500K
test (Seen Targets): 5k
test (New Targets): 5k
github: https://github.com/LAMDASZ-ML/Self-BackTracking
Countdown-Task-GOLDzetagpt-cot-countdown-game-20kcountdown-verl-aha-moment
Countdown Dataset for Verl - DeepSeek-R1 "Aha Moment" Reproduction
This dataset is prepared for reproducing DeepSeek-R1's "aha moment" using the Verl framework. It uses the Countdown task from Jiayi-Pan/Countdown-Tasks-3to4.
Overview
The "aha moment" was first observed in DeepSeek-R1-Zero when training on the Countdown task — the model learns to allocate more thinking time and re-evaluate its initial approach rather than rushing to an answer.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sdfffafag3/countdown-verl-aha-moment.countdownQwen2.5-3B-countdown-level4-1epochs-4rollouts-1024max-length-reasoning-traces-rollout-sftcountdown_problemsCountdowncountdown-numbers-3-8
Countdown Numbers Game Dataset
This dataset contains configurations and solutions for variations of the Countdown numbers game. Each example comprises a sequence of numbers, a target number, the computed solution (closest value), the arithmetic expression that achieves that value, the difference between the target and the computed value, and the final Countdown score.
HuggingFace Download Links
Dataset Variant
Dataset Name
Download
Random… See the full description on the dataset page: https://huggingface.co/datasets/alexjackson17/countdown-numbers-3-8.countdown-numbers-3-8-nz
Countdown Numbers Game Dataset
This dataset contains configurations and solutions for variations of the Countdown numbers game. Each example comprises a sequence of numbers, a target number, the computed solution (closest value), the arithmetic expression that achieves that value, the difference between the target and the computed value, and the final Countdown score.
HuggingFace Download Links
Dataset Variant
Dataset Name
Download
Random… See the full description on the dataset page: https://huggingface.co/datasets/alexjackson17/countdown-numbers-3-8-nz.countdown-sftcountdown_3_groupscountdown-numbers-6-gr
Countdown Numbers Game Dataset
This dataset contains configurations and solutions for variations of the Countdown numbers game. Each example comprises a sequence of numbers, a target number, the computed solution (closest value), the arithmetic expression that achieves that value, the difference between the target and the computed value, and the final Countdown score.
HuggingFace Download Links
Dataset Variant
Dataset Name
Download
Random… See the full description on the dataset page: https://huggingface.co/datasets/alexjackson17/countdown-numbers-6-gr.RAW_DATA-countdown3args-Qwen2.5-1.5B-Instructcountdown-dataset
ES Heterogeneity Countdown
Countdown arithmetic data used for Evolution Strategies experiments under
heterogeneous data allocation.
Dataset splits
train: approximately 3.79 million synthetically generated, solvable, and
deduplicated Countdown problems.
test: 2,000 held-out Countdown problems from the original evaluation set.
Each example contains:
id: example identifier
numbers: input numbers that must each be used exactly once
target: desired arithmetic result… See the full description on the dataset page: https://huggingface.co/datasets/es-heterogeneity/countdown-dataset.countdown-arithmetic-training-pool
Countdown arithmetic training pool
Arithmetic puzzles of the Countdown kind: a handful of source numbers, a target, and the job of
writing an expression over the four operations that reaches the target, using each source number
at most once and not having to use them all. A set generated for this pool and three public
datasets read at the pinned revisions named below, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/countdown-arithmetic-training-pool.llm-cot-countdown-game-20kcountdown-rlvr
Countdown RLVR
Qwen3-4B의 검증 가능한 추론 학습에 사용하는 Countdown 데이터셋입니다.
주어진 숫자를 각각 한 번 사용하여 목표값을 만드는 수식을 생성합니다.
1. 데이터 구성
분할
개수
숫자 개수
목표값
SHA-256
train
1,024
4
10~100
aa7abb6242d8ada72e55a6d8d0917e3618473ddf2f0f880b288814394b231131
validation
128
4
10~100
b06a1be3604d637aa19bd61af57aadf98fbbffcb8ef4db8d477fe8535c617497
test
256
4
10~100
416c02076321875cccfeed19f742e56048269b4b9d24112f6a2bee82ba301815
demo.jsonl에는 검증 흐름을 확인하는 숫자 3개 문제를 둡니다.… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/countdown-rlvr.Countdown-CoT-20k
Rethinking Generalization in Reasoning SFT
This repository contains datasets associated with the paper "Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability".
The research investigates the factors influencing cross-domain generalization in Large Language Models (LLMs) during reasoning-focused supervised fine-tuning (SFT) with long chain-of-thought (CoT) data.
Key Findings
Optimization Dynamics: Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/jasonrqh/Countdown-CoT-20k.9_8_25__countdown_4arg__sft_data_mp_reflection_ckpt_chunk_89_8_25__letter_countdown_4o__sft_data_mp_reflection_ckpt_chunk_5countdown_fresh_heldout_1024
Countdown Fresh Held-Out 1024
This dataset contains 1,024 fresh held-out Countdown arithmetic problems for evaluating language models on the 3-to-4 number Countdown task.
Each problem provides a target integer and a list of 3 or 4 numbers. A model must construct an arithmetic expression using each number at most once and the basic operations +, -, *, and / to equal the target.
Files
standard/test.parquet: prompts for standard no-tool evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/MinghuiXu/countdown_fresh_heldout_1024.countdown-es-grpo-0.1
COUNTDOWN Dataset for ES vs GRPO Comparison
Dataset Description
This dataset contains prepared splits of the COUNTDOWN task for comparing Evolution Strategies (ES) and Group Relative Policy Optimization (GRPO) methods for LLM fine-tuning.
Dataset Statistics
Training split: 10.0% of available data
Training samples: 200
Validation samples: 1,800
Test samples: 200 (reserved for final evaluation)
Data Format
Each example contains:
data: The input… See the full description on the dataset page: https://huggingface.co/datasets/alphaXiv/countdown-es-grpo-0.1.countdown_level_9countdown_datasetcountdown-4dataset__countdown2arg__qwen2.5-1.5b-I__BoN__altered__convos__entropy__base_modelcountdown_level_7countdown-full
COUNTDOWN Dataset for ES vs GRPO Comparison
Dataset Description
Full Countdown dataset (2100 train + 100 test samples) for mathematical reasoning and arithmetic expression generation
Dataset Statistics
Training split: 100.0% of available data
Training samples: 2,100
Validation samples: 0
Test samples: 100 (reserved for final evaluation)
Data Format
Each example contains:
- data: The input prompt/question
- answer: Ground truth answer
-… See the full description on the dataset page: https://huggingface.co/datasets/alphaXiv/countdown-full.countdown-es-grpo-0.4
COUNTDOWN Dataset for ES vs GRPO Comparison
Dataset Description
This dataset contains prepared splits of the COUNTDOWN task for comparing Evolution Strategies (ES) and Group Relative Policy Optimization (GRPO) methods for LLM fine-tuning.
Dataset Statistics
Training split: 40.0% of available data
Training samples: 800
Validation samples: 1,200
Test samples: 200 (reserved for final evaluation)
Data Format
Each example contains:
data: The input… See the full description on the dataset page: https://huggingface.co/datasets/alphaXiv/countdown-es-grpo-0.4.
