reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 5.6 Errors 0 Details… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c.
Upload dataset
Upload README.md with huggingface_hub
Upload dataset
initial commit
