datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime-solution-hint-v6-deepscaler-respgenDeepScaleR_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.aime-solution-hint-v6-deepscaler-respgen__0_115aime-solution-hint-v6-deepscaler-respgen__115_230DeepScaleR-Qwen3-1.7B-0-40kaime-solution-hint-v6-deepscaler-respgen__805_919aime-solution-hint-v6-deepscaler-respgen__230_345aime-solution-hint-v6-deepscaler-respgen__690_805aime-solution-hint-v6-deepscaler-respgen__345_460aime-solution-hint-v6-deepscaler-respgen__460_575aime-solution-hint-v6-deepscaler-respgen__575_690deepscaler_prepare_logp_inputDeepScaleR-Qwen3-1.7B-rl-wholeDeepScaleR-Qwen3-1.7B-23k-classifiedDeepScaleR-Qwen3-1.7B-0-40k-not-all-correctdeepscaler_prepare_logp_input_small_3DeepScaleR-Qwen3-1.7B-22-40kCrosscoder-Qwen2.5-1.5B-vs-DeepScaleR-1.5B_max_activating_examplesSee Files and versions for pickled dictionaries and database versions of of max activating examples organized per available layer, as well as dataframes of available features.
qwen3-8b-base-deepscaler-rollouts
Qwen3-8B-Base rollouts on DeepScaleR, with verifier labels
Frozen on-policy rollouts collected for a two-branch generative-critic study. One
actor, sampled once; every downstream experiment reuses this exact batch.
Generation
actor
Qwen/Qwen3-8B-Base (chat template, enable_thinking=False)
prompts
agentica-org/DeepScaleR-Preview-Dataset, 10,000 sampled (seed 0)
samples per prompt
8
temperature / top-p
0.8 / 0.95
max new tokens
6,144 (context 8… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/qwen3-8b-base-deepscaler-rollouts.deepscaler_prepare_logp_input_smallDeepScaleR-1.5B-Preview_eval_5554
mlfoundations-dev/DeepScaleR-1.5B-Preview_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
41.3
83.0
84.4
29.4
38.7
29.0
33.7
11.6
12.0
13.6
20.3
30.3
21.7
AIME24
Average Accuracy: 41.33% ± 1.26%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepScaleR-1.5B-Preview_eval_5554.deepscaler-Qwen3-8B-Base-4096-n-16DeepScaleR-Qwen3-1.7B-codebook
DeepScaleR-Qwen3-1.7B strategy codebook
116 abstract, problem-independent problem-solving strategies ("codes")
inductively open-coded from Claude's worked solutions to the
zjhhhh/DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded dataset.
Columns
code_id — int, ordered by usage frequency (most-used first).
name — slug (e.g. proof-by-contradiction, pigeonhole, modular-arithmetic).
definition — one-sentence, problem-independent statement of the technique.
example — a… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/DeepScaleR-Qwen3-1.7B-codebook.deepscaler-teacher-sft-vllm-official-40k-clean-v2
DeepScaleR Teacher40k Clean v2
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
minimum official reward: 1.0
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe
Counts
raw examples: 40300
kept examples: 21727
train examples: 21292
val… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v2.DeepScaleR-Preview-Dataset.20000.ancestral.128.Qwen2.5-1.5B-Instructdeepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.deepscaler-Qwen3-1.7B-Base-4096-n-16deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 32768
maximum response chars: 200000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual.omnimath-solution-hint-v6-deepscaler-respgendeepscaler
