reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a
t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking. Games: FrozenLake-v0. Results: 1/1 wins, 0 errors. Dataset Info Rows: 1 Columns: 9 Columns Column Type Description game_id Value('string') Unique identifier for this game episode (env_id + episode number) env_id Value('string') TextArena environment ID (e.g. FrozenLake-v0) model… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a.
t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a
TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking. Games: FrozenLake-v0. Results: 1/1 wins, 0 errors.
Dataset Info
- Rows: 1
- Columns: 9
Columns
Generation Parameters
{
"script_name": "03_arena_eval.py",
"model": "together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking",
"hyperparameters": {
"temperature": 0.7,
"top_p": 0.95,
"max_tokens": 4096,
"enable_thinking": true,
"max_rounds": 30
},
"input_datasets": [],
"description": "TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking. Games: FrozenLake-v0. Results: 1/1 wins, 0 errors.",
"custom_metadata": {
"experiment_name": "t1_synthesize_knowledge_improvement",
"stage": "arena_evaluation",
"run_tag": "d3048f3a",
"system_prompt": null,
"game_configs": [
{
"env_id": "FrozenLake-v0",
"num_episodes": 1,
"num_players": 1
}
]
}
}Usage
from datasets import load_dataset
dataset = load_dataset("reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a", split="train")
print(f"Loaded {len(dataset)} rows")This dataset is tracked in [reasoning-degeneration-dev/PROJECT-MANIFEST](https://huggingface.co/datasets/reasoning-degeneration-dev/PROJECT-MANIFEST)
