reasoning-failures
gepa-rlm-exp-failures_only-20260219-191545
gepa-rlm-exp-failures_only-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: failures_only | Last updated: 2026-02-19 21:50 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
failures_only
46.67%
30.67%
1,150,570
$0.0000
5996s
Learning Curves
Experiment Config
{
"script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-failures_only-20260219-191545.qwen2_5_reasoning_failures
Reasoning and Logic Failure Cases in Qwen2.5-1.5B
Diagnostic dataset of reasoning errors in a small base language model
Technical challenge: Blind Spots of Frontier Models by Fatima Institute for Global AI Research
Overview
This dataset documents systematic reasoning failures observed while evaluating the base language model Qwen/Qwen2.5-1.5B.
The dataset records cases where the model produces confident but incorrect answers to questions requiring:… See the full description on the dataset page: https://huggingface.co/datasets/ALb78/qwen2_5_reasoning_failures.
