abstract-reasoning
gepa-abstract-reasoning-vanilla
gepa-abstract-reasoning-vanilla
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: abstract_reasoning | Last updated: 2026-02-27 01:21 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_vanilla_k3
vanilla
3
abstract_reasoning
31.11%
N/A
55,714
$0.1350
2185s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-abstract-reasoning-vanilla.Abstract_of_Structural_Critique_of_Reasoning
Structural Critique of AI Reasoning Systems — 2025 Comprehensive Review
This repository hosts the condensed 60-page review synthesizing twelve interconnected research papers developed during 2025.
The review presents a unified structural critique of contemporary AI reasoning architectures, demonstrating that their failures arise not from scale or data limitations, but from fundamental architectural misalignments.
📘 Overview
This review reframes all twelve research… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/Abstract_of_Structural_Critique_of_Reasoning.gepa-abstract-reasoning-rlm
gepa-abstract-reasoning-rlm
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: abstract_reasoning | Last updated: 2026-02-27 01:23 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
abstract_reasoning
26.67%
N/A
233,490
$0.0000
2279s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-abstract-reasoning-rlm.gepa-abstract-reasoning-rlm-k10
gepa-abstract-reasoning-rlm-k10
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: abstract_reasoning | Last updated: 2026-02-27 04:34 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k10
rlm
10
abstract_reasoning
28.89%
N/A
469,548
$0.0000
2698s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-abstract-reasoning-rlm-k10.gepa-abstract-reasoning-rlm-k6AbstractReasoning
