CoolFace
Datasetpublic

demfier/reviewertoo-iclr2025-reviews

ReviewerToo — Generated Reviews on ICLR 2025 Reviews generated by ReviewerToo over the full ICLR 2025 submission pool (~11.6k papers). Each paper is reviewed by 11 LLM reviewer personas (monolithic reviews) and synthesized into a single composite metareview with an accept/reject decision. Generated with vllm serving openai/gpt-oss-120b, reasoning-effort=high. Two normalized parquet tables: papers — one row per paper (11,612): metadata, ground-truth program decision, the… See the full description on the dataset page: https://huggingface.co/datasets/demfier/reviewertoo-iclr2025-reviews.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes52downloads
Dataset Card

ReviewerToo — Generated Reviews on ICLR 2025

Reviews generated by ReviewerToo over the full ICLR 2025 submission pool (~11.6k papers). Each paper is reviewed by 11 LLM reviewer personas (monolithic reviews) and synthesized into a single composite metareview with an accept/reject decision. Generated with vllm serving openai/gpt-oss-120b, reasoning-effort=high.

Two normalized parquet tables:

  • —`papers` — one row per paper (11,612): metadata, ground-truth program decision, the ReviewerToo metareviewer decision, and the composite metareview text.
  • —`persona_reviews` — one row per (paper, persona) (127,723): each persona's decision and full review text.

papers fields

fieldtypedescription
paper_idstringOpenReview forum id
yearint2025
title, abstract, primary_areastringsubmission metadata
keywordslist[str]author keywords
arxiv_idstringarXiv id when matched
ground_truthstringreference program decision (raw)
ground_truth_binarystringaccept / reject
reviewer_scoreslisthuman reviewer ratings
metareviewer_decisionstringReviewerToo metareview decision (raw)
metareviewer_decision_binarystringaccept / reject
metareview_textstringcomposite metareview (final synthesis)

persona_reviews fields

fieldtypedescription
paper_idstringOpenReview forum id
personastringreviewer persona key
persona_labelstringhuman-readable persona name
decisionstringpersona's accept/reject recommendation (raw)
decision_binarystringaccept / reject
scorefloatnumeric rating when the persona emitted one — sparse: most personas give a categorical decision rather than a number
review_textstringfull monolithic review (markdown)

Personas (11): Impact (bengio), Visionary (hinton), Fairness (lecun), Probability (pal), Big-picture, Theorist, Pedagogical, Empiricist, Default, Pragmatist, Reproducibility.

Reproducing the metric

Binary accuracy = compare metareviewer_decision_binary to ground_truth_binary on rows where both are set:

  • —metareviewer binary accuracy (full pool, N ≈ 11,585): 71.4%

Notes

  • —Model: openai/gpt-oss-120b via vLLM, reasoning-effort=high; composite metareview synthesized over the 11 monolithic persona reviews.
  • —Ground-truth decisions and ratings come from OpenReview; rejected/withdrawn papers may lack a decision, but their generated reviews are still included.