datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gepa-rlm-exp-algorithmic-20260219-191545
gepa-rlm-exp-algorithmic-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-19 21:34 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
algorithmic
48.89%
26.67%
932,391
$0.0000
5209s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-algorithmic-20260219-191545.gepa-exp-algorithmic-vanilla-20260220-083526
gepa-exp-algorithmic-vanilla-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-20 21:20 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_vanilla_k3
vanilla
3
algorithmic
35.56%
46.67%
77,464
$0.2182
3126s
Learning Curves
Experiment Config
{
"script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-algorithmic-vanilla-20260220-083526.gepa-exp-algorithmic-rlm-20260220-083526
gepa-exp-algorithmic-rlm-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-20 21:25 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
algorithmic
37.78%
44.00%
400,201
$0.0000
3483s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-algorithmic-rlm-20260220-083526.algorithmic-sft-full-eval-v3
algorithmic-sft-full-eval-v3
Aggregate eval results: 16 models × 3 splits (test/harder/OOD). v3 re-evaluation at MAX_TOKENS=32768. Bootstrap 95% CI.
Dataset Info
Rows: 50
Columns: 14
Columns
Column
Type
Description
eval_name
Value('string')
Filename-derived eval identifier
domain
Value('string')
Task domain (countdown, formal_logic, long_arithmetic, cellular_automata, conlang_morphology)
variant
Value('string')
Model variant (algorithm name or… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/algorithmic-sft-full-eval-v3.
