datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MOSAIC-Refactoring
Agentic Pull Request Dataset
Dataset Overview
The dataset contains 4,910,698 Pull Requests in total, consisting of 4,392,818 agent-authored PRs from 10 agents and 517,880 human-authored PRs. The agent-authored PRs come from Claude, Codegen, Codex, Copilot, Cosine, Cursor, Devin, Jules, Junie, and OpenHands. A summary of the dataset is presented below.
Cohort
Pull Requests
Merged Pull Requests
Repositories
Sum of Additions
Sum of Deletions
Humans
517880… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-Refactoring.swe-bench-multi-file-refactoring-sft-dpo-2026
💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents.
📊 Dataset Architecture & Highlights
Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.MOSAIC-Refactoring-copy
Post-Processed Pull Request Dataset
Dataset Overview
The dataset contains a total of 2393 Pull Requests from OpenHands. It also includes additional activity metadata such as repository state snapshots and modified-file records. A summary of the dataset is presented below.
Cohort
Number of PRs
Number of Merged PRs
Unique Repositories
Sum of Additions
Sum of Deletions
OpenHands
2393
1737
667
14186956
1732609
Total
2393
1737
667
14186956
1732609… See the full description on the dataset page: https://huggingface.co/datasets/inaesh-joshi/MOSAIC-Refactoring-copy.Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations
Refactor-Dialogue-1.4k — Multi-turn Refactoring Conversations
Synthetic dataset for fine-tuning coding-focused LLMs on multi-turn
refactoring dialogues. Generated with
Dataset Generator —
an open-source pipeline for building high-quality training data.
Overview
1,414 multi-turn conversations across 3 refactoring categories. Each example
is a 4-message dialogue: user pastes real code → assistant refactors with a
short explanation → user follows up with a constraint or… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations.erlang-refactoringThis repository contains the dataset used for training a proof-of-concept method for refactoring nonidiomatic Erlang code.
This is the implementation of the presented method in the following paper: Balázs Szalontai, Péter Bereczky and Dániel Horpácsi,
"Deep Learning-Based Refactoring with Formally Verified Training Data", Infocommunications Journal, Special Issue on Applied Informatics, 2023, pp. 2-8,
https://doi.org/10.36244/ICJ.2023.5.1
The implementation of the method can be accessed on… See the full description on the dataset page: https://huggingface.co/datasets/szalontaib/erlang-refactoring.repro-codetaste-can-llms-generate-human-level-code-refactorings-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations-Reasoning
Refactor-Dialogue-1.4k-Reasoning — Multi-turn Refactoring Conversations with <think> Reasoning
Synthetic dataset for fine-tuning reasoning-style coding LLMs on
multi-turn refactoring dialogues. Every assistant turn carries a
first-person <think>...</think> internal monologue before the actual
response — DeepSeek-R1 / Qwen3-thinking convention, broadest trainer
compatibility out of the box.
Generated with
Dataset Generator —
an open-source pipeline for building high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations-Reasoning.LLM-Refactoring-Research
