bullets
Datasets
All datasets matching “bullets”nq-item-id-llm-bullets-hardneg-few_shot-v4bullet-swebench-verified
Bullet on SWE-bench Verified — 479/500 = 95.8%
Results for the Bullet coding agent on all 500 instances of
SWE-bench Verified, graded by the official swebench.harness.run_evaluation
scorer. Every instance was attempted and graded; there are no empty patches.
479 / 500 resolved = 95.8%
Run with the Bullet harness on gpt-5.6-sol at high reasoning effort, one attempt per instance.
By repository
repository
resolved
django
223/231
96.5%
sympy
73/75
97.3%… See the full description on the dataset page: https://huggingface.co/datasets/daviddata1/bullet-swebench-verified.exp_8_3_granularity_bullets_test25exp_8_3_granularity_bulletsexp_8_3_granularity_bullets_test5icd11-medgemma-27b-bullets
