datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLMscore-ICLR-OpenReview
LLMscore-ICLR-OpenReview
This dataset is the released original dataset for the paper Position: Peer
Review Should Be Calibrated via LLM Scoring by Zijin Chen, Lesui Yu, Xiaofei
Liao, Hai Jin, and Qinbin Li. The paper has been accepted to the ICML 2026
Position Track.
Its concrete purpose is peer review analysis: the dataset is meant for
studying how paper-review rationales, numeric ratings, LLM-derived anchor
scores, and review-score residuals interact in scientific peer… See the full description on the dataset page: https://huggingface.co/datasets/Wutaghost/LLMscore-ICLR-OpenReview.iclr-rejected-papers-with-code-1k
Rejected ICLR Papers with Reviews and Code
This dataset contains 1,000 rejected ICLR submissions from 2018–2026. Each row
has the OpenReview submission metadata and reviews, the rejected submission PDF,
and a commit-pinned archive of a matched public GitHub repository.
This collection was built directly from OpenReview. It does not use a
third-party ICLR review dataset.
Project repository: TheAppliedScientist
Contents
1,000 unique rejected OpenReview submissions… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-rejected-papers-with-code-1k.benchname-bug-localization
🥷 BenchName (Bug localization)
This is the benchmark for the Bug localization task as part of the
🥷 BenchName benchmark.
The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug.
The dataset provides all the required components for evaluation of bug localization approaches in… See the full description on the dataset page: https://huggingface.co/datasets/anon-iclr-submission/benchname-bug-localization.iclr-papers-with-code-1k
ICLR Papers with Accessible Code
A dataset of 1,051 papers from ICLR (2020-2026) with verified code repositories and complete peer reviews from all reviewers.
Dataset Summary
This dataset contains rejected and borderline-accepted papers from ICLR (International Conference on Learning Representations) with accessible code and full peer review text.
Contents:
1,051 papers total
3,900 reviews (average 3.71 per paper)
944 rejected (90%) + 107 poster-tier accepted… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-papers-with-code-1k.iclr2026-lm-logprobs
LM Log-Probabilities for Value Bias Analysis
Next-token log-probability distributions from 12 language models across 54 prompts, used in the paper:
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A.F. Thompson, Elle, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, Tsvetomira Dumbalska (ICLR 2026)
Part of the Oxford-HIPlab collection for this paper.
Dataset description
Each CSV contains the full next-token log-probability… See the full description on the dataset page: https://huggingface.co/datasets/Oxford-HIPlab/iclr2026-lm-logprobs.ICLR_Peer_Reviews
ICLR Peer Reviews Dataset
This dataset contains peer reviews from ICLR (International Conference on Learning Representations), processed and standardized.
Dataset Structure
title: Title of the paper
abstract: Abstract of the paper
full_text: Full text of the paper (if available)
review: The peer review text
source: Source conference/year (e.g., iclr2020)
review_src: Original source identifier
year: Year of the conference (extracted from source)
overall_score: Computed… See the full description on the dataset page: https://huggingface.co/datasets/JerMa88/ICLR_Peer_Reviews.ICLR-2021-Accepted-Papers
ICLR 2021 International Conference on Learning Representations 2021 Accepted Paper Meta Info Dataset
This dataset is collect from the ICLR 2021 OpenReview website (https://openreview.net/group?id=ICLR.cc/2021/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2021). For researchers who are interested in doing analysis of ICLR 2021 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2021-Accepted-Papers.reviewertoo-iclr2025-reviews
ReviewerToo — Generated Reviews on ICLR 2025
Reviews generated by ReviewerToo over the full ICLR 2025 submission pool
(~11.6k papers). Each paper is reviewed by 11 LLM reviewer personas (monolithic
reviews) and synthesized into a single composite metareview with an accept/reject
decision. Generated with vllm serving openai/gpt-oss-120b, reasoning-effort=high.
Two normalized parquet tables:
papers — one row per paper (11,612): metadata, ground-truth program
decision, the… See the full description on the dataset page: https://huggingface.co/datasets/demfier/reviewertoo-iclr2025-reviews.goodpoint-iclr
GoodPoin-ICLR
Citation
@inproceedings{
anonymous2026goodpoint,
title={GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses},
author={Anonymous},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=u7u3hVx1Io}
}
benchname-module-summarization
🥷 BenchName (Module summarization)
This is the benchmark for Module summarization task as part of the
🥷 BenchName benchmark.
The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects.
The model is required to generate such description, given the relevant context code and the intent behind the documentation.
All the repositories are published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause… See the full description on the dataset page: https://huggingface.co/datasets/anon-iclr-submission/benchname-module-summarization.reviewertoo-iclr2026-reviews
ReviewerToo — Generated Reviews on ICLR 2026
Reviews generated by ReviewerToo over the full ICLR 2026 submission pool.
Each paper is reviewed by 14 LLM reviewer personas (monolithic reviews) and then
synthesized into a single composite metareview with an accept/reject decision.
Generated with vllm serving openai/gpt-oss-120b, reasoning-effort=high.
The dataset ships two normalized parquet tables (the cleaned final outputs) plus
the complete raw run archive (every intermediate… See the full description on the dataset page: https://huggingface.co/datasets/demfier/reviewertoo-iclr2026-reviews.
