Anonymous1-afk-ops/rubric-grounded-faithfulness-eval
Rubric-Grounded Faithfulness Evaluation Resource This repository hosts the anonymized evaluation resource accompanying the NeurIPS 2026 Evaluations and Datasets submission: From Scores to Checks: Rubric-Grounded Faithfulness Evaluation for AI-Generated Images The release contains the derived assets behind the paper's main claims: full-gold AIGCIQA2023 rubric labels, evidence-point and reviewed-counterfactual diagnostic subsets, a relabeled 2400-image T2I-CompBench human-eval… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1-afk-ops/rubric-grounded-faithfulness-eval.
Rubric-Grounded Faithfulness Evaluation Resource
This repository hosts the anonymized evaluation resource accompanying the NeurIPS 2026 Evaluations and Datasets submission:
- From Scores to Checks: Rubric-Grounded Faithfulness Evaluation for AI-Generated Images
The release contains the derived assets behind the paper's main claims: full-gold AIGCIQA2023 rubric labels, evidence-point and reviewed-counterfactual diagnostic subsets, a relabeled 2400-image T2I-CompBench human-eval stress test, and released metric artifacts. It does not redistribute source benchmark images or model weights.
Included assets
data/: derived annotation CSVs and split filesexternal_eval/reference_predictions/: released reference-prediction CSVs used in the paperresults/: released metric and bootstrap artifactsEVALUATION_CARD.md: structured resource documentationcroissant.json: Croissant metadata for this resourceLICENSES.md: upstream license and terms summary
Not included
- source benchmark images from AIGCIQA2023 or T2I-CompBench
- pretrained model weights
- provider credentials
- training or inference code
Intended use
This resource supports:
- full-gold in-domain analysis on AIGCIQA2023
- evidence-grounding and counterfactual diagnostics
- cross-dataset transfer evaluation on the relabeled T2I-CompBench human-eval set
- re-evaluation of released external reference predictions
- inspection of released metric and bootstrap summaries
Upstream images
Users must obtain the original benchmark image assets separately from the upstream benchmark releases and join them to the released CSVs by sample_id and image_name.
Documentation
Start with:
EVALUATION_CARD.mddata/README.mdcroissant.jsonLICENSES.md
Code
The executable code is intentionally hosted separately for the Code URL field. Reviewers should use the OpenReview Code URL together with this dataset/resource bundle.
