CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MoreThought /DeepSWEGym2 Dataset Description This dataset is a filtered and deduplicated version of a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 214.19kb, a total uncompressed size of 17.56GB, and a total of 85974 examples. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2.texttext-generation100K<n<1M1 likes2.3k downloads15d agoHugging Face02MoreThought /DeepSWEGym2-Ultra Dataset Description This dataset is an EXTREMELY filtered and deduplicated version of DeepSWE-Gym2-Edu, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 8.27MB, a total uncompressed size of 8.27GB, and a total of 1000 examples. Dataset Details Curated by: MoreThought Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Ultra.texttext-generationn<1K3 likes1.3k downloads15d agoHugging Face03datacurve /deep-swegated DeepSWE DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks drawn from active open-source repositories. The benchmark includes 113 tasks across TypeScript, Go, Python, JavaScript, and Rust, with isolated environments and program-based verifiers. Task format DeepSWE tasks use the Harbor task format: task.toml Metadata: repository, base commit, language, prebuilt image, resource limits… See the full description on the dataset page: https://huggingface.co/datasets/datacurve/deep-swe.tabularn<1K76 likes1.2k downloads4mo agoHugging Face04MoreThought /DeepSWEGym2-Edu Dataset Description This dataset is a filtered and deduplicated version of a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 291.74kb, a total uncompressed size of 14.93GB, and a total of 53649 examples. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Edu.texttext-generation10K<n<100K2 likes983 downloads16d agoHugging Face05luolc /deep-swe-1-1-materialized DeepSWE 1.1 — materialized A tabular materialization of DeepSWE v1.1 — Datacurve's 113-task benchmark for coding agents — repackaged from datacurve-ai/deep-swe into one parquet row per task. This is a third-party repack for tooling convenience, not an official Datacurve release. Source commit: see manifest.json (source_commit) — every file is carried over unmodified into columns. Integrity: manifest.json records the parquet's sha256 and a per-task content hash (sha256 over each… See the full description on the dataset page: https://huggingface.co/datasets/luolc/deep-swe-1-1-materialized.tabularn<1K0 likes866 downloads28d agoHugging Face06MoreThought /DeepSWEGym2-Full Dataset Description This dataset is a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 169.08kb, a total uncompressed size of 19.80GB, and a total of 122791 examples. Dataset Details Curated by: MoreThought Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Full.texttext-generation100K<n<1M1 likes825 downloads15d agoHugging Face07MoreThought /DeepSWEGym-Edu Dataset Description This dataset is a heavily filtered version of all the SWE-bench/SWE-smith-lang datasets (expect php) merged together (originally 88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with very complex/long code problems in the original datasets, having an average row size of 101.81kb, a total uncompressed size of 4.29GB, and 42529 examples total.… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Edu.texttext-generation10K<n<100K2 likes691 downloads16d agoHugging Face08MoreThought /DeepSWEGym-Full Dataset Description This dataset is a merged version of all the SWE-bench/SWE-smith-lang datasets (88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 85.4kb, a total uncompressed size of 7.53GB, and 88130 examples total. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Full.texttext-generation10K<n<100K1 likes676 downloads16d agoHugging Face09SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-2.8Ktext1K<n<10K8 likes538 downloads1y agoHugging Face10r2e-edits /deepswe-verifier-2582-v1tabular1K<n<10K0 likes514 downloads1y agoHugging Face11LocalLLaMA /deepswe-mini deepswe-mini 16 of the 113 tasks in DeepSWE v1.1, picked so that running just these ranks models the same way the full benchmark does. DeepSWE is a good benchmark and an expensive one. Every task is a long-horizon feature request in its own container, and a full pass takes close to two days of agent time run one task at a time. If you are comparing models, agent harnesses or prompts, and the differences you care about are more than a few points, these 16 tasks give you the same… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/deepswe-mini.tabularn<1K4 likes497 downloads6d agoHugging Face12SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-Samplingtextn<1K0 likes394 downloads9mo agoHugging Face13tangchen-ai /birdcode-deepswe-k1d BirdCode on DeepSWE — k=1, single attempt, no web tools ⚠️ Reading the metric correctly: The summary card's "Average f2p 0.89" is the test-case-level pass fraction (f2p_passed/f2p_total, averaged per task) — it is NOT the official DeepSWE leaderboard metric. The official binary score is the reward field (1 only when all F2P and P2P tests pass): 60/113 = 0.531. Per-trial reward values are visible in each trial's rewards block below. Evaluation of BirdCode (a from-scratch… See the full description on the dataset page: https://huggingface.co/datasets/tangchen-ai/birdcode-deepswe-k1d.tabularn<1K0 likes328 downloads6d agoHugging Face14r2e-edits /deepswe-verifier-merged-with-regression-with-filenamestabular1K<n<10K0 likes234 downloads1y agoHugging Face15r2e-edits /deepswe-verifier-2582-v2tabular1K<n<10K0 likes201 downloads1y agoHugging Face16r2e-edits /deepswe-verifier-merged-with-regressiontabular1K<n<10K0 likes165 downloads1y agoHugging Face17r2e-edits /deepswe-verifier-merged-with-regression-with-filenames-cleantabular1K<n<10K0 likes130 downloads1y agoHugging Face18r2e-edits /deepswe-verifier-3582-exitreason-agent-priority-v1tabular1K<n<10K0 likes112 downloads1y agoHugging Face19r2e-edits /deepswe-swebv-eval-n16-verifier-v1tabular1K<n<10K0 likes110 downloads1y agoHugging Face20r2e-edits /deepswe-verifier-merged-with-regression-with-filenames-and-rewards-v2tabular1K<n<10K0 likes90 downloads1y agoHugging Face21r2e-edits /deepswe-verifier-only-matching-pairs-v1tabular1K<n<10K0 likes87 downloads1y agoHugging Face22r2e-edits /deepswe-verifier-merged-with-regression-with-filenames-and-rewardstabular1K<n<10K1 likes61 downloads1y agoHugging Face23tarsur385 /deepswe-prm-embeddings-8k DeepSWE PRM evaluation embeddings (Qwen3-8B, 8k) Qwen3-8B last-token-pooled embeddings (4096-d) of every step of the DeepSWE evaluation rollouts (mini-swe-agent, Claude-Opus-5, 4 rollouts per task): 44,409 steps, 449 rollouts, 113 tasks. Embedded with preprocessing/deepswe/embed_shard.py at max_model_len 8192. column description trajectory_id rollout id step_idx step order within the rollout task_id DeepSWE task model, config policy model / rollout config… See the full description on the dataset page: https://huggingface.co/datasets/tarsur385/deepswe-prm-embeddings-8k.tabularfeature-extraction10K<n<100K0 likes59 downloads22h agoHugging Face24r2e-edits /deepswe-verifier-exitreason-agent-priority-v1tabular1K<n<10K0 likes54 downloads1y agoHugging Face25r2e-edits /deepswe-swebv-eval-n16-verifier-v1-with-regressiontabular1K<n<10K0 likes49 downloads1y agoHugging Face26DCAgent2 /dev_set_v2_DeepSWE_Preview_20260501_185747textn<1K0 likes42 downloads5mo agoHugging Face27tarsur385 /deepswe-prm-train-embeddings-8k DeepSWE PRM training embeddings (Qwen3-8B, 8k) Frozen Qwen3-8B last-token-pooled embeddings (4096-d, float16) of every step of the DeepSWE training-pool rollouts: 405,919 steps · 4,701 trajectories · 113 tasks. This is the data the released DeepSWE PRM heads (tarsur385/deepswe-prm-heads-8k) were fine-tuned on. Embedded with preprocessing/deepswe/embed_shard.py at max_model_len 8192: the state is the chat-templated step context truncated to its last 8191 tokens; the action is the… See the full description on the dataset page: https://huggingface.co/datasets/tarsur385/deepswe-prm-train-embeddings-8k.tabularfeature-extraction100K<n<1M0 likes33 downloads22h agoHugging Face28r2e-edits /deepswe-verifier-debug-condense-v1tabularn<1K0 likes29 downloads1y agoHugging Face29DCAgent2 /dev_set_v2_rl__r2egym_deepswe_fp8_terminus_2_32b_20260406_202950textn<1K0 likes26 downloads6mo agoHugging Face30DCAgent2 /swebench_verified_random_100_folders_DeepSWE_Preview_20260501_185848textn<1K0 likes26 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.