laion/delphi-1e23-25b-stageE-rl-eval-artifacts
delphi-1e23 (25B) Stage-E RL — raw evalchemy eval artifacts Raw evalchemy (lm-eval v0.4.12) outputs for the delphi-1e23 25B Stage-E RL sweep (marin issue #6279). For each of 5 models (SFT wc50m baseline + 4 RL cells D1–D4) and 2 tasks: <TASK>_<MODEL>_results.json — aggregate metrics + full run config (accuracy for MATH500; exact_match flexible/strict for gsm8k). <TASK>_<MODEL>_samples.jsonl — per-example: problem, gold, model_output, extracted answer, correctness. MATH500 =… See the full description on the dataset page: https://huggingface.co/datasets/laion/delphi-1e23-25b-stageE-rl-eval-artifacts.
073
