datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-8b-base-deepscaler-rollouts
Qwen3-8B-Base rollouts on DeepScaleR, with verifier labels
Frozen on-policy rollouts collected for a two-branch generative-critic study. One
actor, sampled once; every downstream experiment reuses this exact batch.
Generation
actor
Qwen/Qwen3-8B-Base (chat template, enable_thinking=False)
prompts
agentica-org/DeepScaleR-Preview-Dataset, 10,000 sampled (seed 0)
samples per prompt
8
temperature / top-p
0.8 / 0.95
max new tokens
6,144 (context 8… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/qwen3-8b-base-deepscaler-rollouts.deepscaler-teacher-sft-vllm-official-40k-clean-v2
DeepScaleR Teacher40k Clean v2
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
minimum official reward: 1.0
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe
Counts
raw examples: 40300
kept examples: 21727
train examples: 21292
val… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v2.deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 32768
maximum response chars: 200000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual.DeepScaleR-Qwen3-1.7B-2k-strategy-error-200
DeepScaleR Qwen3 1.7B 2K strategy errors
This dataset contains 200 distinct questions selected from
zjhhhh/DeepScaleR-Qwen3-1.7B-2k-agreed-regraded-le5-coded
at revision 8b6e0f481bced00132c95fb631745d4992fa19fd.
Each row has one manually selected model response whose main failure is a
strategy error relative to the source row's code_hint: the response does not
materially use the hint's core route, substitutes another strategy, or omits a
decisive hinted stage in favor of an… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/DeepScaleR-Qwen3-1.7B-2k-strategy-error-200.
