abhid1234/flywheel-rl-benchmark
flywheel — a benchmark for RL on agent trajectories A small, controlled benchmark for studying whether a coding agent can improve from its own graded failures — reinforcement learning at the context layer (the policy update is a durable lesson carried in context, not a weight change) — together with the baseline results from running it live on Daytona sandboxes with a real coding agent. Companion to github.com/abhid1234/flywheel. The full method and honest write-up: FINDINGS.… See the full description on the dataset page: https://huggingface.co/datasets/abhid1234/flywheel-rl-benchmark.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face