datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.code-review-bench
Code Review Bench
A paired online-offline benchmark for AI code review.
Splits
online — Stratified sample of 1,135 bot-reviewed PRs, scraped from open-source Github repositories and scored by the online benchmark (15 tools, Feb–Apr 2026).
offline — 136 expert-curated golden issues across 50 PRs (5 repositories).
Provenance
The offline golden issues extend the 50-PR benchmark originally created by Greptile (2025) and refined by Augment (2025). Our… See the full description on the dataset page: https://huggingface.co/datasets/code-review-bench/code-review-bench.code-review-lab
CodeReview laboratory changes
Synthetic Python before/after pairs and unified diffs for the CodeReview change-scoped secure-review demo. Seed 24.
Organization dataset and collection are public. Live Gradio will be alirezaaminzadeh/code-review and the organization card AriaAICompany/code-review after the daily Space-creation cap resets (scripts/publish.py). Runnable Space source is stored in demo/. Collection: Aria AI — Cybersecurity.
This is fixture data (level 1). The snippets… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/code-review-lab.code-reviewA Scrape of the codereview stack exchange, good for high quality code
code-review
CODE_REVIEW
A preference dataset for CODE_REVIEW, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally code)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits
80/10/10 train /… See the full description on the dataset page: https://huggingface.co/datasets/316usman/code-review.codereview-bench
CodeReview-Bench
A benchmark for evaluating models on two code review tasks, curated from ronantakizawa/github-codereview.
Tasks
1. Code Editing
Given code and a reviewer comment, apply the requested change.
Input: before_code, reviewer_comment, language, diff_context
Target: after_code
from datasets import load_dataset
ds = load_dataset("ronantakizawa/codereview-bench", "code-editing")
example = ds["test"][0]
prompt = f"""Apply the following review comment… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/codereview-bench.reasoning-sft-github-codereview
reasoning-sft-github-codereview
Converted version of ronantakizawa/github-codereview, filtered to 76,689 high-quality rows (quality_score >= 0.75, excluding none comment type).
Nothing fancy, just reformatted the columns into a standard messages format for SFT/reasoning training. No content was modified or regenerated.
Format
Each row has three columns:
input — list of dicts with role and content (system prompt + user turn containing the reviewer comment and original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-github-codereview.code-review-dpo-3k
Code Review DPO Pairs (3K)
DPO preference pairs for training LLMs to produce specific, actionable, educational code reviews.
Dataset Description
3,000 preference pairs across 4 programming languages:
Language
Examples
Python
~64%
JavaScript
~12%
TypeScript
~12%
Go
~12%
8 review scenarios covering real-world code quality issues:
SQL injection & security vulnerabilities
XSS via innerHTML
Hardcoded credentials
Resource leaks (unclosed… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-review-dpo-3k.code-review-python-autotrain
Python Code Review Dataset
Filtered and formatted version of ronantakizawa/github-codereview for fine-tuning code review models.
Dataset Summary
This dataset contains Python code snippets with corresponding review comments, formatted as conversations for instruction tuning.
Splits
Split
Samples
train
~40,000
validation
~800
test
~800
Format
Each sample contains a messages column with conversation format:
{
"messages": [… See the full description on the dataset page: https://huggingface.co/datasets/PrathamKotian26/code-review-python-autotrain.code-review-findings-samples
Code Review Findings Samples
Curated synthetic examples for evaluating automated code review pipelines — especially the
AI Code Reviewer MCP stack built with
Qwen3.6-27B.
Each row contains a short code snippet, the analysis type, and a structured JSON output that
matches the review contract used by ImTamsi/qwen3.6-27b-code-reviewer.
Dataset structure
Column
Description
id
Stable sample identifier
analysis_type
review, bugs, security, performance… See the full description on the dataset page: https://huggingface.co/datasets/ImTamsi/code-review-findings-samples.Code-Review-Assistant
Dataset Card for Code Review Assistant Training Dataset
Dataset Description
Overview
This is the training split of the Code Review Assistant Dataset - a comprehensive synthetic dataset designed for fine-tuning AI models in Python code review, security analysis, and code quality assessment.
Dataset Summary
Curated by: Alen Philip
Language: English (with Python code examples)
License: cc-by-nc-4.0
Total Examples: 13,670
Purpose: Training data for code… See the full description on the dataset page: https://huggingface.co/datasets/alenphilip/Code-Review-Assistant.Code-Review-Assistant-Eval
Dataset Card for Code Review Assistant Evaluation Dataset
Dataset Description
Overview
This is the evaluation split of the Code Review Assistant Dataset - a held-out set for validating and benchmarking models trained on the training dataset. Contains diverse Python code review examples for comprehensive model evaluation.
Dataset Summary
Curated by: Alen Philip
Language: English (with Python code examples)
License: cc-by-nc-4.0
Total Examples: 1,726… See the full description on the dataset page: https://huggingface.co/datasets/alenphilip/Code-Review-Assistant-Eval.
