datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rrma-lean4-agent-traces
RRMA Lean 4 Agent Traces
416 multi-agent Lean 4 proof search traces across two Erdős problems, three model tiers, and four difficulty rungs.
v2 (2026-06-10) — label + format correction. The original upload had two defects:
(1) messages was a JSON string, not an array; (2) reward was set to 1.0 if the
text SCORE=1.0 appeared anywhere in the conversation — including the worker prompt
("repeat until SCORE=1.0") and file reads of the oracle script, so almost every trace was
labeled… See the full description on the dataset page: https://huggingface.co/datasets/vincentoh/rrma-lean4-agent-traces.erdos741ii-lean4-opus-traces
Erdős #741(ii) — Lean 4 Opus Agent Traces
39 Claude Opus agent sessions attempting Erdős problem #741(ii) in Lean 4 (G1 rung: build the proof from an NL construction description).
v2 (2026-06-10) — label correction. The original upload labeled all 39 traces
reward=1.0; the labeler matched the text SCORE=1.0 anywhere in the conversation,
including the worker prompt. Rewards are now anchored to genuine oracle output
(line-anchored SCORE= adjacent to… See the full description on the dataset page: https://huggingface.co/datasets/vincentoh/erdos741ii-lean4-opus-traces.
