CoolFace
Datasetpublic

reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061

t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 1/10 (10.0%) Avg turns 1.4 Errors 4 Details… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes17downloads
3 commits on main
d354bda7mo ago

Upload README.md with huggingface_hub

zsprague
27a2e0d7mo ago

Upload dataset

zsprague
80c62087mo ago

initial commit

zsprague