CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlexCuadron /SWE-Bench-Verified-O1-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.textquestion-answeringn<1K7 likes6.1k downloads2y agoHugging Face02AlexCuadron /SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.textquestion-answeringn<1K4 likes1.9k downloads2y agoHugging Face03swe-qa /SWE-QA-Benchmark SWE-QA Benchmark A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories. Dataset Summary Total Questions: 720 Repositories: 15 Format: JSONL (JSON Lines) Fields: question, answer Repository Coverage Each repository contains 48 questions: astropy conan django flask matplotlib pylint pytest reflex requests scikit-learn sphinx sqlfluff streamlink sympy xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.textquestion-answering1K<n<10K5 likes1.1k downloads2mo agoHugging Face04TIGER-Lab /SWE-QA-Pro-Bench SWE-QA-Pro Bench (A Repository-level QA Benchmark Built from Diverse Long-tail Repositories) 💻 GitHub | 📖 Paper | 🤗 SWE-QA-Pro 📢 News 🚀 [2026-5-19] The evaluation code is released on GitHub. 🔥 [2026-3-23] SWE-QA-Pro Bench is publicly released! The model and code will be released soon. Introduction SWE-QA-Pro Bench is a repository-level question answering dataset designed to evaluate whether models can perform grounded, agentic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-QA-Pro-Bench.textquestion-answeringn<1K5 likes678 downloads4mo agoHugging Face05Anonym01048 /SWE-Bench-Pro-interactive-issue-qa SWE-Bench Pro / Interactive / Issue+QA Companion data release for the anonymous paper "Opinion: Coding-Agent Benchmarks Should Match Their Users' Task Flows" (SWE-TaskFlow). This dataset contains the QA-augmented trajectories: verifiable questions about repository behavior inserted before, between, or after the split issue turns. Every question ships with its hidden reference answer, an executable golden proof script, and the creation-time proof-execution report. The plain… See the full description on the dataset page: https://huggingface.co/datasets/Anonym01048/SWE-Bench-Pro-interactive-issue-qa.textquestion-answering1K<n<10K0 likes245 downloads26d agoHugging Face06beatsprom /swe-bench-multi-file-refactoring-sft-dpo-2026 💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026) High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents. 📊 Dataset Architecture & Highlights Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.texttext-generationn<1K0 likes218 downloads25d agoHugging Face07syntaxsynth /swe-bench-opus-logs Claude 3 inference SWE-Bench results Contains prompting responses from SWE-bench on these 2 settings: Oracle retrieval BM25 retrieval Each of the subsets contains an additional log_last_line attributes which is the last line from log files generated during evaluation step. Results: Model BM25 Retrieval Resolved (%) Oracle Retrieval Resolved (%) GPT-4* 0 1.74 Claude-2 1.96 4.80 Claude-3 Opus (20240229) 3.24 6.42 Claude-2 and GPT-4 results from SWE-bench… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/swe-bench-opus-logs.textquestion-answering1K<n<10K1 likes217 downloads3y agoHugging Face08GeniusWondering /SWE-QA-Benchmark SWE-QA Benchmark A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories. Dataset Summary Total Questions: 720 Repositories: 15 Format: JSONL (JSON Lines) Fields: question, answer Repository Coverage Each repository contains 48 questions: astropy conan django flask matplotlib pylint pytest reflex requests scikit-learn sphinx sqlfluff streamlink sympy xarray… See the full description on the dataset page: https://huggingface.co/datasets/GeniusWondering/SWE-QA-Benchmark.textquestion-answering1K<n<10K0 likes88 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.