datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gepa-abstract-reasoning-vanilla
gepa-abstract-reasoning-vanilla
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: abstract_reasoning | Last updated: 2026-02-27 01:21 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_vanilla_k3
vanilla
3
abstract_reasoning
31.11%
N/A
55,714
$0.1350
2185s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-abstract-reasoning-vanilla.gepa-abstract-reasoning-rlm
gepa-abstract-reasoning-rlm
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: abstract_reasoning | Last updated: 2026-02-27 01:23 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
abstract_reasoning
26.67%
N/A
233,490
$0.0000
2279s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-abstract-reasoning-rlm.gepa-abstract-reasoning-rlm-k10
gepa-abstract-reasoning-rlm-k10
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: abstract_reasoning | Last updated: 2026-02-27 04:34 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k10
rlm
10
abstract_reasoning
28.89%
N/A
469,548
$0.0000
2698s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-abstract-reasoning-rlm-k10.gepa-abstract-reasoning-rlm-k6AbstractReasoning
