datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-bench_Verified_With_AnnotationsSWE-bench_Verified_SmallSWE-bump-benchSWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920
SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency
Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence.
Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.SWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923
Fixed Solo350 u355 + MiniMax-M2.7: orchestration cost study
Best observed cost tradeoff: compact coordinator decisions plus soft review at the existing hard limit (at most 12 worker turns). M2.7 metered token cost falls 59.3%, while mean solved tasks decrease from 90.00 to 87.67/150. Accuracy equivalence was not established.
This closed study contains 6 designs and 16 complete independent runs on the same 150 tasks (2400 scored task/run pairs), each with an independent audit.… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923.
