reasoning-degeneration-dev/gepa-rlm-exp-failures_only-20260219-191545
gepa-rlm-exp-failures_only-20260219-191545 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: failures_only | Last updated: 2026-02-19 21:50 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_rlm_k20 rlm 20 failures_only 46.67% 30.67% 1,150,570 $0.0000 5996s Learning Curves Experiment Config { "script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-failures_only-20260219-191545.
Add configs block to README for dataset viewer
Upload figures/accuracy_vs_tokens.png with huggingface_hub
Upload README.md with huggingface_hub
Upload figures/learning_curve.png with huggingface_hub
Upload figures/accuracy_vs_tokens.png with huggingface_hub
Upload figures/val_score_vs_k.png with huggingface_hub
Upload figures/tokens_vs_k.png with huggingface_hub
Upload figures/accuracy_vs_k.png with huggingface_hub
Upload dataset
Upload dataset
Upload dataset
Upload dataset
Upload state.json with huggingface_hub
Upload figures/learning_curve_live.png with huggingface_hub
Upload dataset
Upload state.json with huggingface_hub
Upload figures/learning_curve_live.png with huggingface_hub
Upload dataset
Upload state.json with huggingface_hub
Upload figures/learning_curve_live.png with huggingface_hub
Upload dataset
Upload state.json with huggingface_hub
Upload figures/learning_curve_live.png with huggingface_hub
Upload dataset
Upload state.json with huggingface_hub
Upload figures/learning_curve_live.png with huggingface_hub
Upload dataset
Upload state.json with huggingface_hub
Upload figures/learning_curve_live.png with huggingface_hub
initial commit
