datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
africa-seychelles-seychelles-economic-social-environmental-health-education-bfdda80b
Seychelles - Economic, Social, Environmental, Health, Education, Development and Energy | Africa (Seychelles official open data)
55,970 rows - 1 Africa country - 1960-2025 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Seychelles as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-seychelles-seychelles-economic-social-environmental-health-education-bfdda80b.wingdings-terminal-bench-2.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260220
wingdings-terminal-bench-2.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260220
Harbor evaluation on terminal-bench@2.0: 0/1 resolved (0.0%), 0 errors
Dataset Info
Rows: 1
Columns: 26
Columns
Column
Type
Description
instance_id
Value('string')
Task identifier (e.g. astropy__astropy-12907)
reward
Value('float64')
Verifier reward (e.g. 0.0 or 1.0)
resolved
Value('bool')
Whether the task was resolved (reward > 0)
agent
Value('string')
Agent… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/wingdings-terminal-bench-2.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260220.wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260222
wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260222
Harbor evaluation on swebench-verified@1.0: 3/10 resolved (30.0%), 0 errors
Dataset Info
Rows: 10
Columns: 26
Columns
Column
Type
Description
instance_id
Value('string')
Task identifier (e.g. astropy__astropy-12907)
reward
Value('float64')
Verifier reward (e.g. 0.0 or 1.0)
resolved
Value('bool')
Whether the task was resolved (reward > 0)
agent
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260222.wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Thinking-20260223
wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Thinking-20260223
Harbor evaluation on swebench-verified@1.0: 1/10 resolved (10.0%), 8 errors
Dataset Info
Rows: 10
Columns: 26
Columns
Column
Type
Description
instance_id
Value('string')
Task identifier (e.g. astropy__astropy-12907)
reward
Value('float64')
Verifier reward (e.g. 0.0 or 1.0)
resolved
Value('bool')
Whether the task was resolved (reward > 0)
agent
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Thinking-20260223.t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-97d2a5df
t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-97d2a5df
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
28.8
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-97d2a5df.t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-67693bd3
t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-67693bd3
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
7/10 (70.0%)
Avg turns
20.4
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-67693bd3.t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-04958e25
t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-04958e25
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
6.2
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-04958e25.t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-9f99c03d
t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-9f99c03d
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
6.4
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-9f99c03d.t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-6ad757f1
t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-6ad757f1
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
5.1
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-6ad757f1.t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-fea1b895
t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-fea1b895
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
9/10 (90.0%)
Avg turns
30.7
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-fea1b895.t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-e3fa99d5
t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-e3fa99d5
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
9/10 (90.0%)
Avg turns
37.1
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-e3fa99d5.Qwen3-Next-80B-MagpieLM-SFT-Outputs-v0.1-shard1-llama8bt1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-9106273f
t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-9106273f
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
6.6
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-9106273f.t1-strategy-musr-critfirst-qwen3-80b-instruct-musr-cf
t1-strategy-musr-critfirst-qwen3-80b-instruct-musr-cf
Strategy compliance evaluation on MuSR murder mysteries — evidence-matrix variant.
Model was instructed to use Criterion-First Cross-Cutting Analysis (evaluate all suspects per criterion before moving to the next).
Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric.
Results
Metric
Value
pass@1
0.9000
Strategy compliance (mean)
4.80
Strategy compliance (min)
4
Strategy… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-critfirst-qwen3-80b-instruct-musr-cf.t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-0b7b235c
t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-0b7b235c
Wingdings compliance evaluation on TextArena games.
Model must reason using only symbolic characters while playing interactive games.
Compliance is measured on reasoning text only, not action commands.
Results
Metric
Value
Win rate
10/10
Compliance (mean)
0.6986
Errors
0
Details
Parameter
Value
Model
together_ai/Qwen/Qwen3-Next-80B-A3B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-0b7b235c.t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-b0ed22ac
t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-b0ed22ac
Wingdings compliance evaluation on TextArena games.
Model must reason using only symbolic characters while playing interactive games.
Compliance is measured on reasoning text only, not action commands.
Results
Metric
Value
Win rate
10/10
Compliance (mean)
0.2177
Errors
0
Details
Parameter
Value
Model
together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-b0ed22ac.t1-strategy-musr-counterfactual-qwen3-80b-thinking-musr-cfact
t1-strategy-musr-counterfactual-qwen3-80b-thinking-musr-cfact
Strategy compliance evaluation on MuSR murder mysteries — counterfactual-hypothesis variant.
Model was instructed to use Counterfactual Hypothesis Testing with Early Termination (assume each suspect is guilty, test conditions, stop at first contradiction).
Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric.
Results
Metric
Value
pass@1
0.9000
Strategy compliance… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-counterfactual-qwen3-80b-thinking-musr-cfact.t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a
t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a
TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking. Games: FrozenLake-v0. Results: 1/1 wins, 0 errors.
Dataset Info
Rows: 1
Columns: 9
Columns
Column
Type
Description
game_id
Value('string')
Unique identifier for this game episode (env_id + episode number)
env_id
Value('string')
TextArena environment ID (e.g. FrozenLake-v0)
model
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a.t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-f51f5776
t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-f51f5776
TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking. Games: FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0. Results: 10/10 wins, 0 errors.
Dataset Info
Rows: 10
Columns: 9
Columns
Column
Type
Description
game_id
Value('string')
Unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-f51f5776.t1-strategy-musr-counterfactual-qwen3-80b-instruct-musr-cfact
t1-strategy-musr-counterfactual-qwen3-80b-instruct-musr-cfact
Strategy compliance evaluation on MuSR murder mysteries — counterfactual-hypothesis variant.
Model was instructed to use Counterfactual Hypothesis Testing with Early Termination (assume each suspect is guilty, test conditions, stop at first contradiction).
Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric.
Results
Metric
Value
pass@1
0.9000
Strategy compliance… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-counterfactual-qwen3-80b-instruct-musr-cfact.t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base
t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base
Strategy compliance evaluation on MuSR murder mysteries — baseline variant.
Model received standard MuSR cot+ prompt (baseline control). Judge still scores against criterion-first rubric.
Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric.
Results
Metric
Value
pass@1
0.9000
Strategy compliance (mean)
2.00
Strategy compliance (min)
2
Strategy compliance (max)
2
Total… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base.t1-strategy-musr-anti-qwen3-80b-instruct-musr-cf
t1-strategy-musr-anti-qwen3-80b-instruct-musr-cf
Strategy compliance evaluation on MuSR murder mysteries — anti-strategy variant.
Model was explicitly told NOT to use matrix/table approach (anti-strategy control). Judge still scores against criterion-first rubric.
Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric.
Results
Metric
Value
pass@1
0.7000
Strategy compliance (mean)
1.60
Strategy compliance (min)
1
Strategy… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-anti-qwen3-80b-instruct-musr-cf.t1-strategy-musr-baseline-qwen3-80b-instruct-musr-base
t1-strategy-musr-baseline-qwen3-80b-instruct-musr-base
Strategy compliance evaluation on MuSR murder mysteries — baseline variant.
Model received standard MuSR cot+ prompt (baseline control). Judge still scores against criterion-first rubric.
Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric.
Results
Metric
Value
pass@1
0.8000
Strategy compliance (mean)
2.90
Strategy compliance (min)
2
Strategy compliance (max)
4
Total… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-baseline-qwen3-80b-instruct-musr-base.t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196
t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
37.0
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196.eval_cleanslate_dataset_qwen3-coder-80bt1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-26f18f00
t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-26f18f00
TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Instruct. Games: FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0. Results: 10/10 wins, 0 errors.
Dataset Info
Rows: 10
Columns: 9
Columns
Column
Type
Description
game_id
Value('string')
Unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-26f18f00.t1-strategy-musr-critfirst-qwen3-80b-thinking-musr-cf
t1-strategy-musr-critfirst-qwen3-80b-thinking-musr-cf
Strategy compliance evaluation on MuSR murder mysteries — criterion-first variant.
Model was instructed to use Criterion-First Cross-Cutting Analysis (evaluate all suspects per criterion before moving to the next).
Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric.
Results
Metric
Value
pass@1
0.8000
Strategy compliance (mean)
4.70
Strategy compliance (min)
4
Strategy… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-critfirst-qwen3-80b-thinking-musr-cf.t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-0898857d
t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-0898857d
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
7/10 (70.0%)
Avg turns
19.5
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-0898857d.t1-strategy-musr-anti-qwen3-80b-thinking-musr-cf
t1-strategy-musr-anti-qwen3-80b-thinking-musr-cf
Strategy compliance evaluation on MuSR murder mysteries — anti-strategy variant.
Model was explicitly told NOT to use matrix/table approach (anti-strategy control). Judge still scores against criterion-first rubric.
Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric.
Results
Metric
Value
pass@1
0.8000
Strategy compliance (mean)
1.50
Strategy compliance (min)
1
Strategy… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-anti-qwen3-80b-thinking-musr-cf.openml-eval-qwen3-next-80b
