CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlexCuadron /SWE-Bench-Verified-O1-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.textquestion-answeringn<1K7 likes6.2k downloads2y agoHugging Face02AlexCuadron /SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.textquestion-answeringn<1K4 likes1.9k downloads2y agoHugging Face03Lego-X /Lego-RL-SWE-Bench-Verified Lego-RL-SWE-Bench-Verified The 500 SWE-bench Verified instances as ready-to-run harbor RL environments — the exact evaluation set behind every SWE-bench Verified number in LEGO-RL, packaged the same way as the training set Lego-X/Lego-RL-2699 so one trainer reads both. Two parallel views of the same 500 instances: View Path What it is Official SWE-bench records swebench_verified_official_500/ The upstream princeton-nlp/SWE-bench_Verified rows, verbatim Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.texttext-generationn<1K0 likes255 downloads1mo agoHugging Face04daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes153 downloads3mo agoHugging Face05JetBrains-Research /agent-trajectories-swe-bench-test-minus-verified Agent Trajectories: SWE-bench Test \ Verified — Mixed Teachers (gpt-5.2 / gpt-5-mini) Summary Full multi-turn agent trajectories collected from the SWE-bench Test minus Verified split (i.e., SWE-bench Test instances that are not part of SWE-bench Verified). Intended for SFT of agent models on coding tasks. Data Collection Each trajectory was produced by a GT-aware lookahead agent that, at every turn: Sampled a candidate response from both gpt-5.2 and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/agent-trajectories-swe-bench-test-minus-verified.tabulartext-generation1K<n<10K0 likes151 downloads6mo agoHugging Face06OpenAgentLab /SWE-Bench_Verified_ABS SWE-Bench_Verified_ABS A dataset of 500 software engineering instances derived from SWE-bench, extended with model-generated test patches for evaluation. GitHub: OpenAgentEval/SWE-ABS Dataset Description Each instance corresponds to a real GitHub issue and pull request. The dataset includes the original SWE-bench fields. test_patch is replaced with a model-generated test patch, and the original is preserved as original_test_patch. Fields Fields inherited from… See the full description on the dataset page: https://huggingface.co/datasets/OpenAgentLab/SWE-Bench_Verified_ABS.texttext-generationn<1K0 likes123 downloads7mo agoHugging Face07CharlieLLL /SWEbench-Verified-eval150-M2.7-solo-selforch-3repeats-w32-20260920 M2.7 Solo and self-orchestration: three independent eval150 runs each All six fresh runs completed the same150 tasks and passed original result/trajectory/task/attempt/fingerprint audits. No previous scores were pooled. Each repeat starts new model processes and cold KV caches after a real telemetry smoke. True failed tasks are retained; infrastructure retries are preserved separately. Mode Repeat1 Repeat2 Repeat3 Mean /150 Sample SD m27-solo 94 98 90 94.00 4.00… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-solo-selforch-3repeats-w32-20260920.text-generation0 likes106 downloads5d agoHugging Face08CharlieLLL /SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920 SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence. Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.tabulartext-generation1K<n<10K0 likes99 downloads5d agoHugging Face09CharlieLLL /SWEbench-Verified-eval150-M2.7-OPD143-appendonly-ablation-2repeats-w32-20260920 M2.7 + mixed-OPD143: coordinator append-only ablation Two new append-only eval150 repeats versus the two existing original-harness repeats (89 and85/150). Same held-out150, models and32-way concurrency; controls were run earlier, not simultaneously. Original controls are reused without rerunning or pooling. Coordinator history Repeat1 Repeat2 Mean accuracy Mean full150 min original 89 85 58.00% 50.77 append_only 80 90 56.67% 110.30 Worker: mixed-OPD… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD143-appendonly-ablation-2repeats-w32-20260920.text-generation0 likes74 downloads5d agoHugging Face10CharlieLLL /SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921 OPD149 bounded evidence7168→3072, latest state at tail: ONE eval150, three 50-task shards 82/150 (54.67%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 69.28 minutes, including queue and any recovery gaps, excluding prior service startup/smoke. Original two-message coordinator/system prompt; only originally visible complete evidence… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921.text-generation0 likes71 downloads5d agoHugging Face11suryadv /devstral-swe-bench-verified Devstral SWE-bench Verified Trajectories This dataset contains 4,000 mini-SWE-agent trajectories generated by mistralai/Devstral-Small-2-24B-Instruct-2512 on a fixed 250-problem training split of SWE-bench Verified. There are 16 independently seeded rollouts per problem. Every row preserves the exact request prefix mini-SWE-agent supplied before Devstral's first generation, the complete ordered rollout, a sanitized final patch, and termination metadata. Official SWE-bench… See the full description on the dataset page: https://huggingface.co/datasets/suryadv/devstral-swe-bench-verified.tabulartext-generation1K<n<10K0 likes68 downloads26d agoHugging Face12CharlieLLL /SWEbench-Verified-eval150-M2.7-new-checkpoints-appendonly-w64-20260920 New coding checkpoints: append-only M2.7 + 9B, 64 episodes Four independent full150 screenings, one repetition each. Historical original-harness and 32-concurrency results are not matched controls; no matched baseline delta or stable gain is claimed. Training arm Checkpoint Resolved /150 Accuracy Full150 min Exact model opd-mix-raw 149 final 93 62.00% 49.92 HF graded350-opd 149 final 80 53.33% 83.08 HF graded124-opd 149 final 86 57.33% 54.82 HF solo350-local63… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-new-checkpoints-appendonly-w64-20260920.text-generation0 likes68 downloads5d agoHugging Face13CharlieLLL /SWEbench-Verified-eval150-M2.7-OPD149-visible-history-v2-2repeats-w64-20260921 OPD149 visible-history v2: two repeats,64 episodes; historical v1 reference Two independent v2 full150 evaluations with OPD149 weights,64 episodes and CPU-quota-aware sandboxes. User requested reusing historical v1 score93/150 rather than a new control run. Historical v1 predates the CPU environment repair and is noncontemporaneous, so this is screening evidence, not a matched causal comparison or proof of non-inferiority. Training arm Checkpoint Resolved /150 Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-visible-history-v2-2repeats-w64-20260921.text-generation0 likes64 downloads5d agoHugging Face14jerry128 /SWE-bench_Verified SWE-bench Verified - Random Subset (100 instances) This is a randomly selected subset of 100 instances from princeton-nlp/SWE-bench_Verified. Dataset Details Source: princeton-nlp/SWE-bench_Verified (split: test) Subset Size: 100 instances Selection Method: Random sampling Random Seed: 42 Created: Automatically generated Instance IDs The following instances are included in this subset: astropy__astropy-13398 astropy__astropy-14508 astropy__astropy-14539… See the full description on the dataset page: https://huggingface.co/datasets/jerry128/SWE-bench_Verified.texttext-generationn<1K0 likes54 downloads10mo agoHugging Face15CharlieLLL /SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921 OPD149 evidence retained, latest state at tail: ONE eval150, three 50-task shards 86/150 (57.33%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 71.76 minutes, including queue and any recovery gaps, excluding prior service startup/smoke. Original two-message coordinator/system prompt; newly exposed evidence events retained, current… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921.text-generation0 likes48 downloads5d agoHugging Face16MemoryAsModality /swebench-verified-kimi-k2p6-traces SWE-bench Verified Kimi K2.6 Reasoning Traces This dataset contains reasoning traces generated on princeton-nlp/SWE-bench_Verified using fireworks_ai/kimi-k2p6-high with a mini-swe-agent based harness. It is intended for research and distillation of software-engineering agents. The repository is published with three configs because each table has a different schema: raw_trajectories: one row per SWE-bench instance with the patch, sanitized result JSON, full trajectory JSON, message… See the full description on the dataset page: https://huggingface.co/datasets/MemoryAsModality/swebench-verified-kimi-k2p6-traces.tabulartext-generation10K<n<100K0 likes41 downloads4mo agoHugging Face17PrimeIntellect /SWE-Bench-Verified-Quick SWE-Bench-Verified-Quick Quick-eval subset of princeton-nlp/SWE-bench_Verified (SWE-bench paper): 468 / 500 Verified instances. Default dataset of the swebench_v1 taskset; also supported by the v0 mini_swe_agent_plus environment. Changes vs upstream Latency subset only — drops the slowest-running instances so a full benchmark pass finishes in under ~30 minutes at reasonable concurrency. No validation semantics; "Verified" in the name is OpenAI's human… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Bench-Verified-Quick.texttext-generationn<1K1 likes31 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.