peer-review
AI-Peer-Review-Detection-Benchmark
Dataset Card for AI Peer Review Detection Benchmark
Dataset Summary
The AI Peer Review Detection Benchmark dataset is the largest dataset to date of paired human- and AI-written peer reviews for identical research papers. It consists of 788,984 reviews generated for 8 years of submissions to two leading AI research conferences: ICLR and NeurIPS. Each AI-generated review is produced using one of five widely-used large language models (LLMs), including GPT-4o, Claude Sonnet… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/AI-Peer-Review-Detection-Benchmark.peerreview-bench
PeerReview Bench
CMU Paper Reviewer:https://prometheus-eval.github.io/cmu-paper-reviewer/
Repository:https://github.com/prometheus-eval/cmu-paper-reviewer
Paper:https://arxiv.org/abs/2605.20668
Point of Contact:seungone@kaist.ac.kr
Expert-annotated review items from scientific papers, organized for three
complementary evaluation tasks. All data in this dataset is intended
for evaluation, not training. All configs reference a shared, deduplicated
file store (submitted_papers)… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/peerreview-bench.openreview-iclr-peer-reviewsopenreview-iclr2025-peer-reviews-RAWai-in-peer-review
Overview
This repository contains more than 50,000 reviews of scientific papers, spanning multiple levels of human-AI collaboration.
Following table summarizes these levels:
Description
Input to LLM
AI-BP: AI-generated with Basic Prompts
Paper + Reviewing guidelines
AI-EP: AI-generated with Elaborate Prompts
Paper + Reviewing guidelines + Conference-issued best practice documents
AI-HI: AI-generated with Human Input
Paper + Reviewing guidelines + Key assessment… See the full description on the dataset page: https://huggingface.co/datasets/rounaksaha12/ai-in-peer-review.ICLR_Peer_Reviews_2026
