Poindexter-Labs/PDL-SWE-Bench
PDL-SWE-Bench PDL-SWE-Bench is an agentic software-engineering benchmark maintained by Poindexter Labs — the SWE sibling of PDL-Bench. Each task drops an agent into an original, internally-authored code repository with an engineering issue written as prose, a passing public test suite, and a fixed token budget. The agent's submitted patch is graded against a held-out acceptance suite it never saw during the episode. Like PDL-Bench, this is an open benchmark (HLE-style): the… See the full description on the dataset page: https://huggingface.co/datasets/Poindexter-Labs/PDL-SWE-Bench.
Rename example_task -> pdl_swe_retry_cap_001 (naming convention); normalize fixture layout; remove .DS_Store artifacts
PDL-SWE-Bench v1.0: 18 tasks (T1-T5), SWE-bench schema + graded-reward extras, dataset card
initial commit
