CoolFace
Datasetpublic

DogukanUrker/PatchEval

PatchEval Pilot 20 PatchEval is a frozen benchmark for agentic coding. Each task starts from the parent of a real Python bug-fix commit whose regression test landed with the fix. The agent receives the reviewed GitHub issue and parent source; scoring is deterministic from hidden fail-to-pass and regression test exit codes. This immutable pilot-20 release contains 20 tasks mined from recent commits and validated in four cells: the hidden regression test fails on the parent; the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/PatchEval.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
1likes104downloads

DogukanUrker/PatchEval · main · files are served by the source, never re-hosted here