cooperdata
cooperdata-sft-midtrain
CooperData — bucketed (SFT / mid-training / discarded)
Unified, coordination-quality-bucketed view of the CooperBench cooperative coding-agent datasets. Each row is one coop pair (two agents each implementing a feature in a shared repo), normalized to a common schema with a source column.
Splits
split
bucket
meaning
rows
sft
A
exemplary coordination workflow worth imitating
2455
midtraining
B
coordination present but thin / one-sided / synthetic
3849… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/cooperdata-sft-midtrain.cooperdata-v3-midtrain-blend
CooperData v3 — Midtraining Blend (Qwen3.5-9B cooperative SWE agents)
All-token midtraining mixture that bridges Qwen/Qwen3.5-9B (instruct) toward the cooperative
multi-agent SWE-coding SFT distribution. One document per row (text, tagged by source) — NOT
packed — so trl.SFTTrainer(packing=False) tokenizes per-doc and the Gated-DeltaNet recurrence
stays per-document. ~390M tokens.
Composition
source
tokens
share
role
web
210.0M
54%
general
math
55.0M… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/cooperdata-v3-midtrain-blend.cooperdata-sft-midtrain-v2
CooperData v2 — contamination-free, coordination-quality-bucketed
Unified view of the CooperBench cooperative coding-agent datasets (+ cooperative-game logs), one row per coop pair, for training a 9B model to be better at CooperBench.
Train/test safety: every pair whose (repo, task_id) is one of the 30 held-out CooperBench benchmark tasks is hard-excluded (X) before bucketing — zero benchmark leakage. team-trajectories and the codex team-coop/cmp-full-team* arms were 100% on… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/cooperdata-sft-midtrain-v2.cooperdata-bridge2x-midtrain-blend
CooperData bridge2x — Midtraining Blend (Qwen3.5-9B cooperative SWE agents)
All-token midtraining mixture (recipe bridge2x) that bridges Qwen/Qwen3.5-9B (instruct)
toward the cooperative multi-agent SWE-coding SFT distribution. One document per row (text,
tagged by source) — NOT packed — so trl.SFTTrainer(packing=False) tokenizes per-doc and the
Gated-DeltaNet recurrence stays per-document. ~200M tokens.
Composition
source
tokens
share
role
coop
120.1M… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/cooperdata-bridge2x-midtrain-blend.cooperdata-bridge-midtrain-blend
CooperData bridge — Midtraining Blend (Qwen3.5-9B cooperative SWE agents)
All-token midtraining mixture (recipe bridge) that bridges Qwen/Qwen3.5-9B (instruct)
toward the cooperative multi-agent SWE-coding SFT distribution. One document per row (text,
tagged by source) — NOT packed — so trl.SFTTrainer(packing=False) tokenizes per-doc and the
Gated-DeltaNet recurrence stays per-document. ~100M tokens.
Composition
source
tokens
share
role
coop
60.0M
60%… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/cooperdata-bridge-midtrain-blend.
