CoolFace
Datasetpublic

reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196

t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 37.0 Errors 0 Details… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
0likes8downloads
Dataset Card

t1-strategy-arena-frozenlake-backwardchaining-togetherai-qwen-qwen3-next-80b-a3b-thin-1d8d6196

Strategy compliance baseline — FrozenLake arena evaluation.

No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review.

Results

MetricValue
Win rate10/10 (100.0%)
Avg turns37.0
Errors0

Details

ParameterValue
Modeltogether_ai/Qwen/Qwen3-Next-80B-A3B-Thinking
ThinkingTrue
Temperature0.7
Max tokens16384
GamesFrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0
Procedural boardsTrue
Seed42
Variantbackward_chaining

Columns

  • —transcript: Full game transcript (JSON array of turn objects)
  • —per_turn_reasoning: Extracted reasoning per turn (JSON array of strings)
  • —reward: Final game reward (1.0 = win, negative = loss)
  • —num_turns: Number of turns played

Usage

python
from datasets import load_dataset
import json

ds = load_dataset("reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196", split="train")
for row in ds:
    reasoning = json.loads(row["per_turn_reasoning"])
    print(f"Game {row['game_id']}: {len(reasoning)} turns, reward={row['reward']}")
    for t, r in enumerate(reasoning):
        if r.strip():
            print(f"  Turn {t}: {r[:200]}...")

Tracked in [reasoning-degeneration-dev/PROJECT-MANIFEST](https://huggingface.co/datasets/reasoning-degeneration-dev/PROJECT-MANIFEST)