CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01electricsheepafrica /africa-seychelles-seychelles-economic-social-environmental-health-education-bfdda80b Seychelles - Economic, Social, Environmental, Health, Education, Development and Energy | Africa (Seychelles official open data) 55,970 rows - 1 Africa country - 1960-2025 - Repackaged by Electric Sheep Africa TL;DR This dataset packages one official CSV resource from Seychelles as ML-ready Parquet. The source file is the provenance boundary; all usable indicators or tabular columns from the resource stay together in this repo. About the source… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-seychelles-seychelles-economic-social-environmental-health-education-bfdda80b.tabulartabular-regression10K<n<100K0 likes334 downloads1mo agoHugging Face02reasoning-degeneration-dev /wingdings-terminal-bench-2.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260220 wingdings-terminal-bench-2.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260220 Harbor evaluation on terminal-bench@2.0: 0/1 resolved (0.0%), 0 errors Dataset Info Rows: 1 Columns: 26 Columns Column Type Description instance_id Value('string') Task identifier (e.g. astropy__astropy-12907) reward Value('float64') Verifier reward (e.g. 0.0 or 1.0) resolved Value('bool') Whether the task was resolved (reward > 0) agent Value('string') Agent… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/wingdings-terminal-bench-2.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260220.tabularn<1K0 likes56 downloads7mo agoHugging Face03reasoning-degeneration-dev /wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260222 wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260222 Harbor evaluation on swebench-verified@1.0: 3/10 resolved (30.0%), 0 errors Dataset Info Rows: 10 Columns: 26 Columns Column Type Description instance_id Value('string') Task identifier (e.g. astropy__astropy-12907) reward Value('float64') Verifier reward (e.g. 0.0 or 1.0) resolved Value('bool') Whether the task was resolved (reward > 0) agent Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Instruct-20260222.tabularn<1K0 likes29 downloads7mo agoHugging Face04reasoning-degeneration-dev /wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Thinking-20260223 wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Thinking-20260223 Harbor evaluation on swebench-verified@1.0: 1/10 resolved (10.0%), 8 errors Dataset Info Rows: 10 Columns: 26 Columns Column Type Description instance_id Value('string') Task identifier (e.g. astropy__astropy-12907) reward Value('float64') Verifier reward (e.g. 0.0 or 1.0) resolved Value('bool') Whether the task was resolved (reward > 0) agent Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/wingdings-swebench-verified-1.0-mini-swe-agent-Qwen3-Next-80B-A3B-Thinking-20260223.tabularn<1K0 likes24 downloads7mo agoHugging Face05reasoning-degeneration-dev /t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-97d2a5df t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-97d2a5df Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 28.8 Errors 0 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-97d2a5df.tabularn<1K0 likes20 downloads7mo agoHugging Face06reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-67693bd3 t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-67693bd3 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 7/10 (70.0%) Avg turns 20.4 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-67693bd3.tabularn<1K0 likes18 downloads7mo agoHugging Face07reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-04958e25 t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-04958e25 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 6.2 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-04958e25.tabularn<1K0 likes17 downloads7mo agoHugging Face08reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-9f99c03d t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-9f99c03d Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 6.4 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-9f99c03d.tabularn<1K0 likes15 downloads7mo agoHugging Face09reasoning-degeneration-dev /t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-6ad757f1 t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-6ad757f1 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 5.1 Errors 0 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-inst-6ad757f1.tabularn<1K0 likes15 downloads7mo agoHugging Face10reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-fea1b895 t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-fea1b895 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 9/10 (90.0%) Avg turns 30.7 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-fea1b895.tabularn<1K0 likes15 downloads7mo agoHugging Face11reasoning-degeneration-dev /t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-e3fa99d5 t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-e3fa99d5 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 9/10 (90.0%) Avg turns 37.1 Errors 0 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-e3fa99d5.tabularn<1K0 likes14 downloads7mo agoHugging Face12yunjae-won /Qwen3-Next-80B-MagpieLM-SFT-Outputs-v0.1-shard1-llama8btabular100K<n<1M0 likes13 downloads1y agoHugging Face13reasoning-degeneration-dev /t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-9106273f t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-9106273f Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 6.6 Errors 0 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-qwen-qwen3-next-80b-a3b-thin-9106273f.tabularn<1K0 likes12 downloads7mo agoHugging Face14reasoning-degeneration-dev /t1-strategy-musr-critfirst-qwen3-80b-instruct-musr-cf t1-strategy-musr-critfirst-qwen3-80b-instruct-musr-cf Strategy compliance evaluation on MuSR murder mysteries — evidence-matrix variant. Model was instructed to use Criterion-First Cross-Cutting Analysis (evaluate all suspects per criterion before moving to the next). Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.9000 Strategy compliance (mean) 4.80 Strategy compliance (min) 4 Strategy… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-critfirst-qwen3-80b-instruct-musr-cf.tabularn<1K0 likes11 downloads7mo agoHugging Face15reasoning-degeneration-dev /t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-0b7b235c t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-0b7b235c Wingdings compliance evaluation on TextArena games. Model must reason using only symbolic characters while playing interactive games. Compliance is measured on reasoning text only, not action commands. Results Metric Value Win rate 10/10 Compliance (mean) 0.6986 Errors 0 Details Parameter Value Model together_ai/Qwen/Qwen3-Next-80B-A3B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-0b7b235c.tabularn<1K0 likes10 downloads7mo agoHugging Face16reasoning-degeneration-dev /t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-b0ed22ac t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-b0ed22ac Wingdings compliance evaluation on TextArena games. Model must reason using only symbolic characters while playing interactive games. Compliance is measured on reasoning text only, not action commands. Results Metric Value Win rate 10/10 Compliance (mean) 0.2177 Errors 0 Details Parameter Value Model together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-b0ed22ac.tabularn<1K0 likes10 downloads7mo agoHugging Face17reasoning-degeneration-dev /t1-strategy-musr-counterfactual-qwen3-80b-thinking-musr-cfact t1-strategy-musr-counterfactual-qwen3-80b-thinking-musr-cfact Strategy compliance evaluation on MuSR murder mysteries — counterfactual-hypothesis variant. Model was instructed to use Counterfactual Hypothesis Testing with Early Termination (assume each suspect is guilty, test conditions, stop at first contradiction). Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.9000 Strategy compliance… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-counterfactual-qwen3-80b-thinking-musr-cfact.tabularn<1K0 likes10 downloads7mo agoHugging Face18reasoning-degeneration-dev /t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking. Games: FrozenLake-v0. Results: 1/1 wins, 0 errors. Dataset Info Rows: 1 Columns: 9 Columns Column Type Description game_id Value('string') Unique identifier for this game episode (env_id + episode number) env_id Value('string') TextArena environment ID (e.g. FrozenLake-v0) model Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-d3048f3a.tabularn<1K0 likes9 downloads7mo agoHugging Face19reasoning-degeneration-dev /t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-f51f5776 t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-f51f5776 TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Thinking. Games: FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0. Results: 10/10 wins, 0 errors. Dataset Info Rows: 10 Columns: 9 Columns Column Type Description game_id Value('string') Unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-thin-f51f5776.tabularn<1K0 likes8 downloads7mo agoHugging Face20reasoning-degeneration-dev /t1-strategy-musr-counterfactual-qwen3-80b-instruct-musr-cfact t1-strategy-musr-counterfactual-qwen3-80b-instruct-musr-cfact Strategy compliance evaluation on MuSR murder mysteries — counterfactual-hypothesis variant. Model was instructed to use Counterfactual Hypothesis Testing with Early Termination (assume each suspect is guilty, test conditions, stop at first contradiction). Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.9000 Strategy compliance… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-counterfactual-qwen3-80b-instruct-musr-cfact.tabularn<1K0 likes8 downloads7mo agoHugging Face21reasoning-degeneration-dev /t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base Strategy compliance evaluation on MuSR murder mysteries — baseline variant. Model received standard MuSR cot+ prompt (baseline control). Judge still scores against criterion-first rubric. Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.9000 Strategy compliance (mean) 2.00 Strategy compliance (min) 2 Strategy compliance (max) 2 Total… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-baseline-qwen3-80b-thinking-musr-base.tabularn<1K0 likes8 downloads7mo agoHugging Face22reasoning-degeneration-dev /t1-strategy-musr-anti-qwen3-80b-instruct-musr-cf t1-strategy-musr-anti-qwen3-80b-instruct-musr-cf Strategy compliance evaluation on MuSR murder mysteries — anti-strategy variant. Model was explicitly told NOT to use matrix/table approach (anti-strategy control). Judge still scores against criterion-first rubric. Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.7000 Strategy compliance (mean) 1.60 Strategy compliance (min) 1 Strategy… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-anti-qwen3-80b-instruct-musr-cf.tabularn<1K0 likes7 downloads7mo agoHugging Face23reasoning-degeneration-dev /t1-strategy-musr-baseline-qwen3-80b-instruct-musr-base t1-strategy-musr-baseline-qwen3-80b-instruct-musr-base Strategy compliance evaluation on MuSR murder mysteries — baseline variant. Model received standard MuSR cot+ prompt (baseline control). Judge still scores against criterion-first rubric. Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.8000 Strategy compliance (mean) 2.90 Strategy compliance (min) 2 Strategy compliance (max) 4 Total… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-baseline-qwen3-80b-instruct-musr-base.tabularn<1K0 likes7 downloads7mo agoHugging Face24reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196 t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 37.0 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-thin-1d8d6196.tabularn<1K0 likes7 downloads7mo agoHugging Face25unlearning-cleanslate /eval_cleanslate_dataset_qwen3-coder-80btabularn<1K0 likes7 downloads7mo agoHugging Face26reasoning-degeneration-dev /t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-26f18f00 t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-26f18f00 TextArena interactive evaluation of together_ai/Qwen/Qwen3-Next-80B-A3B-Instruct. Games: FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0, FrozenLake-v0. Results: 10/10 wins, 0 errors. Dataset Info Rows: 10 Columns: 9 Columns Column Type Description game_id Value('string') Unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-arena-frozenlake-together_ai-qwen-qwen3-next-80b-a3b-inst-26f18f00.tabularn<1K0 likes6 downloads7mo agoHugging Face27reasoning-degeneration-dev /t1-strategy-musr-critfirst-qwen3-80b-thinking-musr-cf t1-strategy-musr-critfirst-qwen3-80b-thinking-musr-cf Strategy compliance evaluation on MuSR murder mysteries — criterion-first variant. Model was instructed to use Criterion-First Cross-Cutting Analysis (evaluate all suspects per criterion before moving to the next). Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.8000 Strategy compliance (mean) 4.70 Strategy compliance (min) 4 Strategy… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-critfirst-qwen3-80b-thinking-musr-cf.tabularn<1K0 likes6 downloads7mo agoHugging Face28reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-0898857d t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-0898857d Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 7/10 (70.0%) Avg turns 19.5 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-qwen-qwen3-next-80b-a3b-inst-0898857d.tabularn<1K0 likes6 downloads7mo agoHugging Face29reasoning-degeneration-dev /t1-strategy-musr-anti-qwen3-80b-thinking-musr-cf t1-strategy-musr-anti-qwen3-80b-thinking-musr-cf Strategy compliance evaluation on MuSR murder mysteries — anti-strategy variant. Model was explicitly told NOT to use matrix/table approach (anti-strategy control). Judge still scores against criterion-first rubric. Compliance is scored by an LLM judge (1-5 Likert) against the Criterion-First rubric. Results Metric Value pass@1 0.8000 Strategy compliance (mean) 1.50 Strategy compliance (min) 1 Strategy… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-musr-anti-qwen3-80b-thinking-musr-cf.tabularn<1K0 likes5 downloads7mo agoHugging Face30asingh15 /openml-eval-qwen3-next-80btabularn<1K0 likes4 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.