datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchname-bug-localization
🥷 BenchName (Bug localization)
This is the benchmark for the Bug localization task as part of the
🥷 BenchName benchmark.
The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug.
The dataset provides all the required components for evaluation of bug localization approaches in… See the full description on the dataset page: https://huggingface.co/datasets/anon-iclr-submission/benchname-bug-localization.iclr-rejected-papers-with-code-1k
Rejected ICLR Papers with Reviews and Code
This dataset contains 1,000 rejected ICLR submissions from 2018–2026. Each row
has the OpenReview submission metadata and reviews, the rejected submission PDF,
and a commit-pinned archive of a matched public GitHub repository.
This collection was built directly from OpenReview. It does not use a
third-party ICLR review dataset.
Project repository: TheAppliedScientist
Contents
1,000 unique rejected OpenReview submissions… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-rejected-papers-with-code-1k.iclr2026-lm-logprobs
LM Log-Probabilities for Value Bias Analysis
Next-token log-probability distributions from 12 language models across 54 prompts, used in the paper:
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A.F. Thompson, Elle, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, Tsvetomira Dumbalska (ICLR 2026)
Part of the Oxford-HIPlab collection for this paper.
Dataset description
Each CSV contains the full next-token log-probability… See the full description on the dataset page: https://huggingface.co/datasets/Oxford-HIPlab/iclr2026-lm-logprobs.ICLR_Peer_Reviews
ICLR Peer Reviews Dataset
This dataset contains peer reviews from ICLR (International Conference on Learning Representations), processed and standardized.
Dataset Structure
title: Title of the paper
abstract: Abstract of the paper
full_text: Full text of the paper (if available)
review: The peer review text
source: Source conference/year (e.g., iclr2020)
review_src: Original source identifier
year: Year of the conference (extracted from source)
overall_score: Computed… See the full description on the dataset page: https://huggingface.co/datasets/JerMa88/ICLR_Peer_Reviews.goodpoint-iclr
GoodPoin-ICLR
Citation
@inproceedings{
anonymous2026goodpoint,
title={GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses},
author={Anonymous},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=u7u3hVx1Io}
}
