CoolFace
Datasetpublic

CooperBench/qwen35-9b-contract-first-coop-random-50

What this is Cooperative two-agent coding dataset: 36 task pairs across 13 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a contract-first prompt variant — agents first agree on a shared interface contract (function signatures, data structures, API boundaries) before independently implementing their respective features. All 36 pairs were successfully evaluated. At a glance Field Value Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-contract-first-coop-random-50.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes22downloads
Dataset Card

What this is

Cooperative two-agent coding dataset: 36 task pairs across 13 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a contract-first prompt variant — agents first agree on a shared interface contract (function signatures, data structures, API boundaries) before independently implementing their respective features. All 36 pairs were successfully evaluated.

At a glance

FieldValue
ModelQwen/Qwen3.5-9B
Agentminisweagent (contract-first prompt)
Settingcoop
Subsetrandom-50 (36 completed)
Repos13
Pairs evaluated36
Both-pass5.6% (2/36)
Per-feature pass20.8% (15/72)
Merge clean rate69.4% (25/36)
Merge conflict rate30.6% (11/36)
Total tokens~20.2M (in+out, from traj files)
OwnerArya Prabhudesai
Date2026-06-02

How it was generated

bash
cooperbench run --setting coop -a mini_swe_agent -c 30 qwen35-9b-contract-first-coop-random-50

Model served via vLLM OpenAI-compatible endpoint (openai/Qwen/Qwen3.5-9B). Contract-first variant: agents negotiate and agree on a shared interface contract before beginning implementation, aiming to reduce integration conflicts.

File layout

  • index.csv — one row per task pair; HF Dataset Viewer entry point
  • qwen35-9b-contract-first-coop-random-50/coop/<repo>/<task_id>/<features>/ — raw per-pair artifacts: result.json, eval.json, agent1_traj.json, agent2_traj.json, agent{1,2}.patch, conversation.json

log_dir column in index.csv points to the per-pair subdirectory.

Schema highlights for mid-training

Filter on: both_passed=true, model, agent_framework.

metadata JSON carries per-agent statuses, steps, merge outcome, per-feature pass — use json.loads(row["metadata"]) without following the pointer.

Note: total_tokens is 0 for this run — token counts are in agent_full_traj.json under messages[*].extra.response.usage (~20.2M total in+out).

Caveats

  • random-50 subset — 36 tasks completed
  • 16.7% LimitsExceeded exits (12/72 agent slots); 1.4% Error (1/72)
  • 30.6% merge conflict rate (11/36)
  • Token fields in result.json are 0; aggregate from agent_full_traj.json if needed

Citation

bibtex
@dataset{qwen35_9b_contract_first_coop_random_50,
  title  = {qwen35-9b-contract-first-coop-random-50},
  author = {Arya Prabhudesai},
  year   = {2026},
  url    = {https://huggingface.co/datasets/CooperBench/qwen35-9b-contract-first-coop-random-50}
}