CooperBench/qwen35-9b-contract-first-coop-random-50
What this is Cooperative two-agent coding dataset: 36 task pairs across 13 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a contract-first prompt variant — agents first agree on a shared interface contract (function signatures, data structures, API boundaries) before independently implementing their respective features. All 36 pairs were successfully evaluated. At a glance Field Value Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-contract-first-coop-random-50.
What this is
Cooperative two-agent coding dataset: 36 task pairs across 13 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a contract-first prompt variant — agents first agree on a shared interface contract (function signatures, data structures, API boundaries) before independently implementing their respective features. All 36 pairs were successfully evaluated.
At a glance
How it was generated
cooperbench run --setting coop -a mini_swe_agent -c 30 qwen35-9b-contract-first-coop-random-50Model served via vLLM OpenAI-compatible endpoint (openai/Qwen/Qwen3.5-9B). Contract-first variant: agents negotiate and agree on a shared interface contract before beginning implementation, aiming to reduce integration conflicts.
File layout
index.csv— one row per task pair; HF Dataset Viewer entry pointqwen35-9b-contract-first-coop-random-50/coop/<repo>/<task_id>/<features>/— raw per-pair artifacts:result.json,eval.json,agent1_traj.json,agent2_traj.json,agent{1,2}.patch,conversation.json
log_dir column in index.csv points to the per-pair subdirectory.
Schema highlights for mid-training
Filter on: both_passed=true, model, agent_framework.
metadata JSON carries per-agent statuses, steps, merge outcome, per-feature pass — use json.loads(row["metadata"]) without following the pointer.
Note: total_tokens is 0 for this run — token counts are in agent_full_traj.json under messages[*].extra.response.usage (~20.2M total in+out).
Caveats
- random-50 subset — 36 tasks completed
- 16.7% LimitsExceeded exits (12/72 agent slots); 1.4% Error (1/72)
- 30.6% merge conflict rate (11/36)
- Token fields in
result.jsonare 0; aggregate fromagent_full_traj.jsonif needed
Citation
@dataset{qwen35_9b_contract_first_coop_random_50,
title = {qwen35-9b-contract-first-coop-random-50},
author = {Arya Prabhudesai},
year = {2026},
url = {https://huggingface.co/datasets/CooperBench/qwen35-9b-contract-first-coop-random-50}
}