CoolFace
Datasetpublic

reasoning-degeneration-dev/gepa-rlm-exp-failures_only-20260219-191545

gepa-rlm-exp-failures_only-20260219-191545 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: failures_only | Last updated: 2026-02-19 21:50 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_rlm_k20 rlm 20 failures_only 46.67% 30.67% 1,150,570 $0.0000 5996s Learning Curves Experiment Config { "script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-failures_only-20260219-191545.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes62downloads

reasoning-degeneration-dev/gepa-rlm-exp-failures_only-20260219-191545 · main · files are served by the source, never re-hosted here