CoolFace
Datasetpublic

Anonymous1-afk-ops/rubric-grounded-faithfulness-eval

Rubric-Grounded Faithfulness Evaluation Resource This repository hosts the anonymized evaluation resource accompanying the NeurIPS 2026 Evaluations and Datasets submission: From Scores to Checks: Rubric-Grounded Faithfulness Evaluation for AI-Generated Images The release contains the derived assets behind the paper's main claims: full-gold AIGCIQA2023 rubric labels, evidence-point and reviewed-counterfactual diagnostic subsets, a relabeled 2400-image T2I-CompBench human-eval… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1-afk-ops/rubric-grounded-faithfulness-eval.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes121downloads
Dataset Card

Rubric-Grounded Faithfulness Evaluation Resource

This repository hosts the anonymized evaluation resource accompanying the NeurIPS 2026 Evaluations and Datasets submission:

  • —From Scores to Checks: Rubric-Grounded Faithfulness Evaluation for AI-Generated Images

The release contains the derived assets behind the paper's main claims: full-gold AIGCIQA2023 rubric labels, evidence-point and reviewed-counterfactual diagnostic subsets, a relabeled 2400-image T2I-CompBench human-eval stress test, and released metric artifacts. It does not redistribute source benchmark images or model weights.

Included assets

  • —data/: derived annotation CSVs and split files
  • —external_eval/reference_predictions/: released reference-prediction CSVs used in the paper
  • —results/: released metric and bootstrap artifacts
  • —EVALUATION_CARD.md: structured resource documentation
  • —croissant.json: Croissant metadata for this resource
  • —LICENSES.md: upstream license and terms summary

Not included

  • —source benchmark images from AIGCIQA2023 or T2I-CompBench
  • —pretrained model weights
  • —provider credentials
  • —training or inference code

Intended use

This resource supports:

  • —full-gold in-domain analysis on AIGCIQA2023
  • —evidence-grounding and counterfactual diagnostics
  • —cross-dataset transfer evaluation on the relabeled T2I-CompBench human-eval set
  • —re-evaluation of released external reference predictions
  • —inspection of released metric and bootstrap summaries

Upstream images

Users must obtain the original benchmark image assets separately from the upstream benchmark releases and join them to the released CSVs by sample_id and image_name.

Documentation

Start with:

  1. 1.EVALUATION_CARD.md
  2. 2.data/README.md
  3. 3.croissant.json
  4. 4.LICENSES.md

Code

The executable code is intentionally hosted separately for the Code URL field. Reviewers should use the OpenReview Code URL together with this dataset/resource bundle.