demfier/reviewertoo-iclr2025-reviews
ReviewerToo — Generated Reviews on ICLR 2025 Reviews generated by ReviewerToo over the full ICLR 2025 submission pool (~11.6k papers). Each paper is reviewed by 11 LLM reviewer personas (monolithic reviews) and synthesized into a single composite metareview with an accept/reject decision. Generated with vllm serving openai/gpt-oss-120b, reasoning-effort=high. Two normalized parquet tables: papers — one row per paper (11,612): metadata, ground-truth program decision, the… See the full description on the dataset page: https://huggingface.co/datasets/demfier/reviewertoo-iclr2025-reviews.
ReviewerToo — Generated Reviews on ICLR 2025
Reviews generated by ReviewerToo over the full ICLR 2025 submission pool (~11.6k papers). Each paper is reviewed by 11 LLM reviewer personas (monolithic reviews) and synthesized into a single composite metareview with an accept/reject decision. Generated with vllm serving openai/gpt-oss-120b, reasoning-effort=high.
Two normalized parquet tables:
- `papers` — one row per paper (11,612): metadata, ground-truth program decision, the ReviewerToo metareviewer decision, and the composite metareview text.
- `persona_reviews` — one row per (paper, persona) (127,723): each persona's decision and full review text.
papers fields
persona_reviews fields
Personas (11): Impact (bengio), Visionary (hinton), Fairness (lecun), Probability (pal), Big-picture, Theorist, Pedagogical, Empiricist, Default, Pragmatist, Reproducibility.
Reproducing the metric
Binary accuracy = compare metareviewer_decision_binary to ground_truth_binary on rows where both are set:
- metareviewer binary accuracy (full pool, N ≈ 11,585): 71.4%
Notes
- Model:
openai/gpt-oss-120bvia vLLM,reasoning-effort=high; composite metareview synthesized over the 11 monolithic persona reviews. - Ground-truth decisions and ratings come from OpenReview; rejected/withdrawn papers may lack a decision, but their generated reviews are still included.
