CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CharlieLLL /SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920 SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence. Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.tabulartext-generation1K<n<10K0 likes98 downloads3d agoHugging Face02CooperBench /qwen9b-coop-claude-code qwen9b-coop-claude-code Two-agent cooperative coding trajectories generated by running CooperBench in coop mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote. The matched solo (single-agent) baseline is at CooperBench/qwen9b-solo-claude-code. Same task corpus, same model, same agent — only the… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code.tabulartext-generationn<1K0 likes77 downloads4mo agoHugging Face03CooperBench /qwen9b-solo-claude-code qwen9b-solo-claude-code Single-agent coding trajectories generated by running CooperBench in solo mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. One agent implements both features in each task. The matched coop (two-agent) version is at CooperBench/qwen9b-coop-claude-code. Same task corpus, same model, same agent — only the coordination differs, so together they isolate the cooperation deficit. At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.tabulartext-generationn<1K0 likes71 downloads4mo agoHugging Face04CooperBench /qwen9b-coop-mini-swe-agent qwen9b-coop-mini-swe-agent Two-agent cooperative coding trajectories generated by running CooperBench in coop mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework. Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote. The matched solo version is at CooperBench/qwen9b-solo-mini-swe-agent. Same task corpus, same model, same agent — only the coordination differs… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-mini-swe-agent.tabulartext-generationn<1K0 likes59 downloads4mo agoHugging Face05CooperBench /qwen9b-solo-mini-swe-agent qwen9b-solo-mini-swe-agent Single-agent coding trajectories generated by running CooperBench in solo mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework. One agent implements both features in each task. The matched coop version is at CooperBench/qwen9b-coop-mini-swe-agent. Same task corpus, same model, same agent — only the coordination differs, so together they isolate the cooperation deficit. At a glance… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-mini-swe-agent.tabulartext-generationn<1K0 likes50 downloads4mo agoHugging Face06sboughorbel /diffing-stats-gemma-2-9b-it-L20-k100-lr1e-04-Crosscodertabular100K<n<1M0 likes46 downloads1y agoHugging Face07CooperBench /qwen35-9b-plan-first-coop What this is Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a plan-first prompt variant — agents are prompted to produce an explicit implementation plan before writing code, then coordinate to reconcile plans before proceeding. Patches are auto-merged after both submit. At a glance Field Value Model Qwen/Qwen3.5-9B Agent mini_swe_agent (plan-first prompt)… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-plan-first-coop.tabularn<1K0 likes46 downloads3mo agoHugging Face08CooperBench /qwen35-9b-git-coop What this is Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting with a shared read-only git remote (--git). Agents coordinate via messaging and git fetch team; patches are auto-merged after both submit. At a glance Field Value Model Qwen/Qwen3.5-9B Agent mini_swe_agent (step_limit=300) Setting coop + git remote Repos 18 Pairs 211 Both-pass 5.7% (12/210… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-git-coop.tabularn<1K0 likes40 downloads3mo agoHugging Face09sboughorbel /diffing-stats-gemma-2-9b-it-L20-mu1.0e-01-lr1e-04-local-shuffling-CrosscoderLosstabular100K<n<1M0 likes37 downloads1y agoHugging Face10sboughorbel /diffing-stats-gemma-2-9b-it-DPO-L20-k100-lr1e-04-dpo-simpo-Crosscodertabular100K<n<1M0 likes31 downloads1y agoHugging Face11CooperBench /qwen35-9b-question-first-coop-random-50 What this is Cooperative two-agent coding dataset: 49 task pairs across 15 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a question-first prompt variant — agents begin by asking each other clarifying questions about their respective features before starting implementation, aiming to surface integration concerns early. All 49 pairs were successfully evaluated. At a glance Field Value Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-question-first-coop-random-50.tabularn<1K0 likes30 downloads3mo agoHugging Face12MaryGonzalez /minor-pool-9b57ec minor-pool-9b57ec Synthetic sensors test data: 55 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/MaryGonzalez/minor-pool-9b57ec.tabularn<1K0 likes30 downloads13d agoHugging Face13CooperBench /qwen35-9b-explore-plan-coop What this is Cooperative two-agent coding dataset: 209 task pairs across 18 repos, generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using an explore-plan prompt variant — agents explore the codebase first, then produce an explicit implementation plan, then share and reconcile plans with their partner before writing code. Patches are auto-merged after both submit. At a glance Field Value Model Qwen/Qwen3.5-9B Agent mini_swe_agent… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-explore-plan-coop.tabularn<1K0 likes26 downloads3mo agoHugging Face14CooperBench /qwen35-9b-milestone-checkins-coop What this is Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a milestone-checkins prompt variant — agents use periodic structured check-ins at predefined milestones to coordinate progress and surface integration conflicts early. Patches are auto-merged after both submit. Coverage caveat: Only 146 of 211 pairs were successfully evaluated (65 had eval errors). The high agent Error rate… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-milestone-checkins-coop.tabularn<1K0 likes22 downloads3mo agoHugging Face15CooperBench /qwen35-9b-contract-first-coop-random-50 What this is Cooperative two-agent coding dataset: 36 task pairs across 13 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a contract-first prompt variant — agents first agree on a shared interface contract (function signatures, data structures, API boundaries) before independently implementing their respective features. All 36 pairs were successfully evaluated. At a glance Field Value Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-contract-first-coop-random-50.tabularn<1K0 likes22 downloads3mo agoHugging Face16CooperBench /qwen9b-coop-claude-code-compressed qwen9b-coop-claude-code-compressed-ak Synthetic compressed cooperative agent trajectories derived from CooperBench/qwen9b-coop-claude-code. Each raw pair (two LLM coding agents on overlapping features in the same repo, communicating via Redis messaging + a team git remote) is condensed into an idealized version: wasted steps dropped, broken submission rituals fixed, missing cooperation events (coop-send/coop-broadcast/coop-recv, git fetch/diff team) inserted where the real pair… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code-compressed.tabularn<1K0 likes21 downloads4mo agoHugging Face17CooperBench /qwen35-9b-late-sync-coop-random-50 What this is Cooperative two-agent coding dataset: 48 task pairs across 15 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a late-sync prompt variant — agents work independently for most of the task and synchronise only at a late stage before submission. Patches are auto-merged after both submit. All 48 pairs were successfully evaluated. At a glance Field Value Model Qwen/Qwen3.5-9B Agent mini_swe_agent… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-late-sync-coop-random-50.tabularn<1K0 likes21 downloads3mo agoHugging Face18sboughorbel /diffing-stats-gemma-2-9b-L20-k100-lr1e-04-base-it-Crosscodertabular100K<n<1M0 likes20 downloads1y agoHugging Face19Chakshu /before-during-after-b703f136-4fb7-4603-9bee-9ca01138bcfdtabular1K<n<10K0 likes18 downloads3y agoHugging Face20CooperBench /qwen35-9b-reasoning-share-coop-random-50 What this is Cooperative two-agent coding dataset: 50 task pairs across 15 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a reasoning-share prompt variant — agents share their internal reasoning and analysis with each other before and during implementation, giving each agent visibility into the other's thought process to improve integration. All 50 pairs were successfully evaluated. At a glance Field Value… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-reasoning-share-coop-random-50.tabularn<1K0 likes17 downloads3mo agoHugging Face21sboughorbel /diffing-stats-gemma-2-9b-L20-k100-lr1e-04-Crosscodertabular100K<n<1M0 likes16 downloads1y agoHugging Face22CooperBench /qwen35-9b-leader-follower-coop What this is Cooperative two-agent coding dataset: 39 task pairs across 14 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a leader-follower prompt variant — one agent is designated as leader and sets the coordination strategy; the other acts as follower and adapts its implementation plan accordingly. All 39 pairs were successfully evaluated. Notable: this variant produced the lowest merge conflict rate (17.9%) of all random-50… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-leader-follower-coop.tabularn<1K0 likes15 downloads3mo agoHugging Face23CooperBench /qwen35-9b-async-coop What this is Cooperative two-agent coding dataset: 50 task pairs across 15 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using an async prompt variant — agents work fully asynchronously without active coordination (no messaging, no synchronisation points). Patches are auto-merged after both submit. All 50 pairs were successfully evaluated. At a glance Field Value Model Qwen/Qwen3.5-9B Agent mini_swe_agent… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-async-coop.tabularn<1K0 likes15 downloads3mo agoHugging Face24Chakshu /equake-a0700949-bff4-49a9-9b97-d81e057a9066tabular1K<n<10K0 likes12 downloads3y agoHugging Face25sboughorbel /diffing-stats-gemma-2-9b-L20-k100-lr1e-04-base-dpo-Crosscodertabular100K<n<1M0 likes11 downloads1y agoHugging Face26sboughorbel /diffing-stats-gemma-2-9b-it-L20-k100-lr1e-04-CCLosstabular100K<n<1M0 likes7 downloads1y agoHugging Face27CooperBench /qwen35-9b-test-impl-split-coop What this is Cooperative two-agent coding dataset: 48 task pairs across 15 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a test-impl-split prompt variant — one agent is assigned the role of writing tests; the other writes the implementation. The two patches cover non-overlapping files, which eliminates merge conflicts entirely. Key finding: This variant achieves a 100% clean merge rate (0 conflicts across all 48 pairs), but a 0%… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-test-impl-split-coop.tabularn<1K0 likes7 downloads3mo agoHugging Face28sboughorbel /diffing-stats-gemma-2-9b-it-SimPO-L20-mu1.0e-01-lr1e-04-CrosscoderLosstabular100K<n<1M0 likes3 downloads1y agoHugging Face29sboughorbel /diffing-stats-gemma-2-9b-it-L20-mu4.0e-02-lr1e-04-CrosscoderLosstabular100K<n<1M0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.