CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ronantakizawa /github-codereview Code Review Dataset A large-scale dataset of the best human-written code reviews from top GitHub repositories. Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response. The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable. This provides a natural signal for training models to: Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.tabulartext-generation100K<n<1M62 likes1.4k downloads7mo agoHugging Face02Dahoas /code-review-instruct-critique-revision Dataset Card for "code-review-instruct-critique-revision" More Information needed text10K<n<100K4 likes216 downloads4y agoHugging Face03ruoyu001 /swebench-codereview-benchmark-v3 SWE-bench Code Review Benchmark v3 This dataset contains 7 benchmark splits for evaluating code review models on the SWE-bench task. Dataset Summary Total instances: 3500 Total resolved: 801 (22.9%) Splits: 7 (3 main + 4 weak models) Version: 3.0.0 Created: 2026-05-03 Splits Split Instances Resolved Resolve Rate Model glm5_500_v3 500 361 72.2% openai/GLM-5-FP8 qwen3_coder_30b_500_v3 500 235 47.0% Qwen/Qwen3-Coder-30B-A3B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/ruoyu001/swebench-codereview-benchmark-v3.text1K<n<10K0 likes148 downloads5mo agoHugging Face04liodon-ai /gemma4-code-review-instruct gemma4-code-review-instruct 197K code review examples — 58K with chain-of-thought <think> reasoning traces. Built to train models that don't just flag issues, but explain their reasoning before delivering a review. Drop-in ready for SFT with any chat model. Why This Dataset Most code review datasets give you diff → comment. This one gives you diff → think → comment for 30% of examples — reasoning traces that show how to analyze a diff before writing the review.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/gemma4-code-review-instruct.texttext-generation100K<n<1M4 likes121 downloads3mo agoHugging Face05code-review-bench /code-review-bench Code Review Bench A paired online-offline benchmark for AI code review. Splits online — Stratified sample of 1,135 bot-reviewed PRs, scraped from open-source Github repositories and scored by the online benchmark (15 tools, Feb–Apr 2026). offline — 136 expert-curated golden issues across 50 PRs (5 repositories). Provenance The offline golden issues extend the 50-PR benchmark originally created by Greptile (2025) and refined by Augment (2025). Our… See the full description on the dataset page: https://huggingface.co/datasets/code-review-bench/code-review-bench.tabulartext-generation1K<n<10K1 likes115 downloads2mo agoHugging Face06sujalgawas /multilang-code-quality-reviewstext10K<n<100K0 likes113 downloads1mo agoHugging Face07Dahoas /code-review-instruct-critique-revision-pythontext1K<n<10K10 likes84 downloads4y agoHugging Face08Dahoas /base_code_review Dataset Card for "base_code_review" More Information needed text10K<n<100K1 likes83 downloads4y agoHugging Face09manishsaini1 /github-codereview-dataset Github-Codereview-Dataset Made with ❤️ using 🦥 Unsloth Studio github-codereview-dataset was generated with Unsloth Recipe Studio. It contains 10,000 generated records. 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("manishsaini1/github-codereview-dataset", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 10,000 📋 Columns: 23 📋 Schema & Statistics… See the full description on the dataset page: https://huggingface.co/datasets/manishsaini1/github-codereview-dataset.tabular10K<n<100K1 likes82 downloads9d agoHugging Face10sungjin-code /amazon-reviews-for-llm-extended Cross-domain sequential recommendation dataset A sequential recommendation dataset drawn from Amazon Reviews 2023, covering Books, CDs_and_Vinyl, Movies_and_TV, Video_Games. Each row of interactions.parquet is one user buying or reviewing one item at one time. Users are sampled so that every one of them is active in all domains, their interactions are ordered chronologically and cut into train/valid/test, and each interaction carries a fixed set of 10 candidate items for ranking… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-for-llm-extended.tabular1K<n<10K0 likes68 downloads29d agoHugging Face11316usman /code-review CODE_REVIEW A preference dataset for CODE_REVIEW, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally code) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits 80/10/10 train /… See the full description on the dataset page: https://huggingface.co/datasets/316usman/code-review.texttext-generation1K<n<10K0 likes62 downloads14d agoHugging Face12reshinthadith /2048_has_code_filtered_base_code_review_python Dataset Card for "2048_has_code_filtered_base_code_review_python" More Information needed text1K<n<10K0 likes56 downloads4y agoHugging Face13mlfoundations-dev /stackexchange_codereviewtext10K<n<100K1 likes54 downloads2y agoHugging Face14ronantakizawa /codereview-bench CodeReview-Bench A benchmark for evaluating models on two code review tasks, curated from ronantakizawa/github-codereview. Tasks 1. Code Editing Given code and a reviewer comment, apply the requested change. Input: before_code, reviewer_comment, language, diff_context Target: after_code from datasets import load_dataset ds = load_dataset("ronantakizawa/codereview-bench", "code-editing") example = ds["test"][0] prompt = f"""Apply the following review comment… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/codereview-bench.texttext-generation100K<n<1M3 likes50 downloads7mo agoHugging Face15DCAgent2 /terminal_bench_2_a1_stackexchange_codereview_20260711_155918text1K<n<10K0 likes50 downloads2mo agoHugging Face16DCAgent /stackexchange-codereview-sandboxes_glm_4.7_traces_jupitertext10K<n<100K0 likes43 downloads6mo agoHugging Face17sujalgawas /big-multilang-code-quality-reviewstext10K<n<100K0 likes42 downloads1mo agoHugging Face18mlfoundations-dev /stackexchange-codereview-sandboxes-traces-terminus-2text1K<n<10K0 likes40 downloads1y agoHugging Face19reshinthadith /2048_has_code_filtered_base_code_review_python_based_on_property Dataset Card for "2048_has_code_filtered_base_code_review_python_based_on_property" More Information needed text1K<n<10K0 likes39 downloads4y agoHugging Face20mlfoundations-dev /b2_code_fasttext_pos_codeforces_neg_codereviewtabular10K<n<100K0 likes38 downloads1y agoHugging Face21sungjin-code /amazon-reviews-for-llm Cross-domain sequential recommendation dataset A sequential recommendation dataset drawn from Amazon Reviews 2023, covering Books, CDs_and_Vinyl, Movies_and_TV, Video_Games. Each row of interactions.parquet is one user buying or reviewing one item at one time. Users are sampled so that every one of them is active in all domains, their interactions are ordered chronologically and cut into train/valid/test, and each interaction carries a fixed set of 10 candidate items for ranking… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-for-llm.tabular1K<n<10K0 likes33 downloads1mo agoHugging Face22mlfoundations-dev /stackexchange-codereview-sandboxestext10K<n<100K0 likes29 downloads1y agoHugging Face23AmanPriyanshu /reasoning-sft-github-codereview reasoning-sft-github-codereview Converted version of ronantakizawa/github-codereview, filtered to 76,689 high-quality rows (quality_score >= 0.75, excluding none comment type). Nothing fancy, just reformatted the columns into a standard messages format for SFT/reasoning training. No content was modified or regenerated. Format Each row has three columns: input — list of dicts with role and content (system prompt + user turn containing the reviewer comment and original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-github-codereview.texttext-generation10K<n<100K0 likes29 downloads7mo agoHugging Face24sungjin-code /amazon-reviews-books-for-llm Books sequential recommendation dataset A sequential recommendation dataset drawn from Amazon Reviews 2023, covering Books. Each row of interactions.parquet is one user buying or reviewing one item at one time. Users are sampled so that every one of them is active in all domains, their interactions are ordered chronologically and cut into train/valid/test, and each interaction carries a fixed set of 10 candidate items for ranking evaluation. Integer user_idx / item_idx columns… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-books-for-llm.tabular10K<n<100K0 likes29 downloads1mo agoHugging Face25Code-TREAT /code_review_generationtext100K<n<1M0 likes27 downloads1y agoHugging Face26PrathamKotian26 /code-review-python-autotrain Python Code Review Dataset Filtered and formatted version of ronantakizawa/github-codereview for fine-tuning code review models. Dataset Summary This dataset contains Python code snippets with corresponding review comments, formatted as conversations for instruction tuning. Splits Split Samples train ~40,000 validation ~800 test ~800 Format Each sample contains a messages column with conversation format: { "messages": [… See the full description on the dataset page: https://huggingface.co/datasets/PrathamKotian26/code-review-python-autotrain.texttext-generation10K<n<100K0 likes23 downloads6mo agoHugging Face27DCAgent /exp_8_1_style_transfer_code_review_comment_test25textn<1K0 likes22 downloads7mo agoHugging Face28DCAgent2 /terminal_bench_2_stackexchange_codereview_sandboxes_traces_terminus_2_overwrite6d215e47textn<1K0 likes22 downloads6mo agoHugging Face29Dahoas /4096_filtered_base_code_review Dataset Card for "4096_filtered_base_code_review" More Information needed text10K<n<100K3 likes21 downloads4y agoHugging Face30DCAgent /exp_8_1_style_transfer_code_review_comment_test5textn<1K0 likes21 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.