t46/atlas-of-judgment
Atlas of Judgment — ICLR Peer-Review Logic Units 1,420,178 atomic units of evaluative logic extracted from public ICLR peer reviews (OpenReview), each decomposed into what was inspected → what was observed → how it was reasoned about → what was concluded, and labeled with a data-induced taxonomy of 12 objects of scrutiny × 12 epistemic standards. This is the dataset behind the Atlas of Judgment (interactive visual atlas). Full reproduction record — every script, model, parameter… See the full description on the dataset page: https://huggingface.co/datasets/t46/atlas-of-judgment.
Atlas of Judgment — ICLR Peer-Review Logic Units
1,420,178 atomic units of evaluative logic extracted from public ICLR peer reviews (OpenReview), each decomposed into what was inspected → what was observed → how it was reasoned about → what was concluded, and labeled with a data-induced taxonomy of 12 objects of scrutiny × 12 epistemic standards.
This is the dataset behind the Atlas of Judgment (interactive visual atlas). Full reproduction record — every script, model, parameter, seed, count, cost, and failure mode — at atlas-of-judgment.pages.dev/method and github.com/t46/atlas-of-judgment.
⚠️ AI-generated content disclosure
Every row is a machine reading of a human review, produced by a two-stage LLM pipeline (DeepSeek memo → Qwen structuring). Units are grounded in the reviewers' public text, but the decomposition, phrasing, and labels are model output and may contain errors or omissions; they are for research reference only. The support_status field records whether a unit is grounded in the reviewer's explicit words or inferred by the memo layer. Known instrument bias: 71.8% of units resolve negative — a property of the extraction instrument as much as of reviewers.
Subsets
The 2026 review-level track reads 98.1% of the venue's official reviews at the finest grain; review_id is the OpenReview note id, so every unit can be held against the original review. The forum-level track reads 50,861 of 51,813 public forums (98.2%) across nine years and adds discussion-phase fields (temporal_position, judgment_change, update_trigger).
taxonomy/taxonomy_v1.json carries the label definitions (12 objects × 12 standards, induced by UMAP+HDBSCAN over a 12k sample, then human-named). taxonomy/rhetoric_labels_analyst.csv carries 600 analyst-labeled units for the six inference forms (NORM / BLOCK / DOUBT / ANCHOR / REACH / WEIGH) used by the atlas's rhetoric classifier.
Fields (both unit tables)
inspected_object/observation/reasoning/judgment— the four-step logic decomposition, in the pipeline's wordsvalence— negative · positive · conditional · uncertain · mixedsuggested_improvement— the concrete fix, when the review offered onesupport_status— explicitly_supported vs memo-inferred groundingconfidence— the structuring model's own confidence in the unitobject_key,object_sim,standard_key,standard_sim— nearest-centroid taxonomy assignment with cosine similarity to the centroid- forum-level only:
year,forum_id,reviewer_key(per-forum anonymous reviewer id),reviewer_role,temporal_position(initial vs post-author-response),judgment_change(maintained / strengthened / weakened / reversed),update_trigger - review-level only:
paper_id,review_id(OpenReview note ids),n_evidence_refs,n_missing_links
Provenance & pipeline
- Raw: public OpenReview API v1/v2, 52,460 ICLR forums, 2018–2026.
- Memos:
deepseek-v4-flash(temp 0.4) writes qualitative-metascience memos — review-level (ICLR 2026) and forum-level (2018–2026) tracks. - Units:
qwen3.7-flashnormalizes memos into the structured units here. - Taxonomy: induced from a 12k-unit sample (bge-small-en-v1.5 embeddings, UMAP n=15 / HDBSCAN 60/10, seed 7), human-named, assigned corpus-wide by nearest centroid.
Decisions, scores, and other outcome metadata were never shown to the extraction pipeline; join them from OpenReview at analysis time if needed.
Analysis artifacts (analysis/)
Every derived dataset behind the Atlas of Judgment figures (atlas-of-judgment.pages.dev), 29 JSON files. Highlights:
elements-all(.raw).json,grounds-all-raw.json,noveltylaw-raw.json,argument-raw-novelty.json,caselaw-raw.json— the element decomposition of every charge (ground → law → remedy): raw k-means clusters with exemplars and per-unit assignments, plus the merged/named layers. The cluster structure is data-driven; the merges and names are readings of exemplars (analyst for novelty; machine-drafted, analyst-reviewed for the rest) — these files exist so the borders can be re-drawn.jurisprudence-data.json— the within-paper tariff of each objection, price-vs-repairability, same-bench/different-law agreement.boilerplate-data.json— nearest-twin similarity of criticisms across papers (measured on the instrument's normalized readings; absolute levels are upper bounds).counterfactual-data.json— panel-redraw simulation (variance components, empirical acceptance curve, per-paper flip probabilities).threads-data.json— reply-tree shapes of all 199k review threads, 2018–2026.- plus the itinerary, court (meta-review), formulary, search-party, tide, nine-year-grammar, interrogative, deliberation, lottery, lifecycle, repertoire, currents, overrule, repair, oracle, and season tables.
All machine readings; hold them against the originals on OpenReview.
License & attribution
Released under CC BY 4.0, matching the upstream license of the source material: reviews and comments on OpenReview are released by their submitters under CC BY 4.0 (OpenReview Terms of Use), and ICLR 2026 submissions carry a CC BY 4.0 license field. Please attribute OpenReview and the ICLR reviewer community (source text) and cite this dataset (machine reading). Reviewer identities are the public anonymous ids; no non-public personal data is included.
@misc{atlas-of-judgment-2026,
title = {Atlas of Judgment: ICLR Peer-Review Logic Units},
author = {Takagi, Shiro},
year = {2026},
url = {https://huggingface.co/datasets/t46/atlas-of-judgment}
}Limitations
- Machine reading, not ground truth: validate before drawing strong claims.
- 1.9% of 2026 reviews and 1.8% of forums failed extraction; missingness is not verified to be topic-random.
- The standard taxonomy is a skeleton induced from ~40% of the pilot sample (61.5% HDBSCAN noise on the reasoning side); per-unit assignment similarity is provided so you can filter.
- Sub-scores and decisions are not included here — join them from OpenReview.
