datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f
t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f.t1-structured-reasoning-user-full-together_ai-qwen-qwen3-next-80b-a3b-thin-a18d7696
t1-structured-reasoning-user-full-together_ai-qwen-qwen3-next-80b-a3b-thin-a18d7696
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-user-full-together_ai-qwen-qwen3-next-80b-a3b-thin-a18d7696.t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297
t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time format… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297.t1-structured-reasoning-full-together_ai-minimaxai-minimax-m2-5-d87b36d3
t1-structured-reasoning-full-together_ai-minimaxai-minimax-m2-5-d87b36d3
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-minimaxai-minimax-m2-5-d87b36d3.t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-thin-11c2eead
t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-thin-11c2eead
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-thin-11c2eead.t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-thin-3fe3254b
t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-thin-3fe3254b
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-thin-3fe3254b.t1-musr-structured-murder-together_ai-qwen-qwen3-next-80b-a3b-inst-4d7e26c9
t1-musr-structured-murder-together_ai-qwen-qwen3-next-80b-a3b-inst-4d7e26c9
Structured slot-filling evaluation for MuSR murder mysteries. Instead of the
original paper's freeform cot+ hint, this uses a structured algorithm that forces
the model to extract facts, slot-fill means/motive/opportunity, and compare suspects.
Results
Metric
Value
pass@1
0.8000
Total problems
10
Structured Prompt (excerpt)
You are solving a murder mystery. You must… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-musr-structured-murder-together_ai-qwen-qwen3-next-80b-a3b-inst-4d7e26c9.t1-structured-reasoning-full-together_ai-zai-org-glm-5-16c46443
t1-structured-reasoning-full-together_ai-zai-org-glm-5-16c46443
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time format… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-zai-org-glm-5-16c46443.m1-1k-tokenized-v3-structured_reasoning-0503t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-inst-a090ce60
t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-inst-a090ce60
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-inst-a090ce60.t1-musr-structured-murder-together_ai-qwen-qwen3-next-80b-a3b-thin-caaf5585
t1-musr-structured-murder-together_ai-qwen-qwen3-next-80b-a3b-thin-caaf5585
Structured slot-filling evaluation for MuSR murder mysteries. Instead of the
original paper's freeform cot+ hint, this uses a structured algorithm that forces
the model to extract facts, slot-fill means/motive/opportunity, and compare suspects.
Comparison to Baseline
Metric
Baseline (cot+)
Structured (slot-fill)
Delta
pass@1
0.8000
0.7000
-0.1000
Results… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-musr-structured-murder-together_ai-qwen-qwen3-next-80b-a3b-thin-caaf5585.t1-structured-reasoning-full-kimi-k2-instruct-0905
t1-structured-reasoning-full-kimi-k2-instruct-0905
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time format instruction
alone… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-kimi-k2-instruct-0905.m1-1k-tokenized-formatted-structured-reasoning-0504gsm8k-structured-reasoningt1-structured-reasoning-user-full-together_ai-qwen-qwen3-next-80b-a3b-inst-64c082d4
t1-structured-reasoning-user-full-together_ai-qwen-qwen3-next-80b-a3b-inst-64c082d4
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-user-full-together_ai-qwen-qwen3-next-80b-a3b-inst-64c082d4.t1-structured-reasoning-full-together_ai-qwen-qwen3-5-397b-a17b-5f29dd2a
t1-structured-reasoning-full-together_ai-qwen-qwen3-5-397b-a17b-5f29dd2a
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-qwen-qwen3-5-397b-a17b-5f29dd2a.t1-structured-reasoning-full-together_ai-qwen-qwen3-235b-a22b-thinkin-7f11a2fe
t1-structured-reasoning-full-together_ai-qwen-qwen3-235b-a22b-thinkin-7f11a2fe
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-qwen-qwen3-235b-a22b-thinkin-7f11a2fe.
