CoolFace
Datasetpublic

WRBench/wrbench-human-annotations

WRBench Human Annotations This dataset contains the human comparison labels used to validate WRBench's automatic evaluation metrics. Version Update: 2026-07-07 We updated the release after rechecking videos that changed during benchmark maintenance. The release now includes: 1,741 clean comparison rows. 4,302 individual human judgments. 585 newly rechecked current-benchmark comparisons, each reviewed by three annotators. Majority-label summaries for the newly… See the full description on the dataset page: https://huggingface.co/datasets/WRBench/wrbench-human-annotations.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes82downloads
Dataset Card

WRBench Human Annotations

This dataset contains the human comparison labels used to validate WRBench's automatic evaluation metrics.

Version Update: 2026-07-07

We updated the release after rechecking videos that changed during benchmark maintenance. The release now includes:

  • 1,741 clean comparison rows.
  • 4,302 individual human judgments.
  • 585 newly rechecked current-benchmark comparisons, each reviewed by three annotators.
  • Majority-label summaries for the newly rechecked subset.

These counts describe this dated human-label release. Moving leaderboard, video, and model counts are reported by the results and videos datasets.

Videos for the current recheck

All 1,170 endpoints in current_review_consensus resolve to the exact bytes shown to annotators. Of these, 897 are in wrbench-videos; the remaining 273 endpoints refer to 106 unique files bundled here under videos/ and indexed by asset_manifest.parquet.

The bundled files are dated human-review evidence. They use the earlier profile-unsuffixed IDs and are not additional leaderboard videos or aliases for the 30-degree and 60-degree benchmark profiles. Join on video_asset_id, whose sha256: value verifies the file bytes.

Configs

  • pairs: clean comparison rows used by the public WRBench release.
  • annotation_to_video: maps each comparison to the corresponding video assets.
  • current_review_consensus: majority labels for the 585 newly rechecked comparisons.
  • current_review_submissions: anonymized per-annotator labels for those 585 comparisons.
  • current_review_status and exclusions: records kept for transparency, but not used as clean comparison labels.
  • review_assets: paths, byte counts, and SHA256 identifiers for the 106 review-only videos bundled with this dataset.

Annotator identities are anonymized in the public files.

Links

  • Videos: https://huggingface.co/datasets/WRBench/wrbench-videos
  • Results: https://huggingface.co/datasets/WRBench/wrbench-results
  • Prompts: https://huggingface.co/datasets/WRBench/wrbench-natural25
  • Leaderboard: https://huggingface.co/spaces/WRBench/wrbench-leaderboard
  • Collection: https://huggingface.co/collections/WRBench/wrbench-current-world-models-lack-a-persistent-state-core
  • Paper: https://arxiv.org/abs/2606.20545
  • Code: https://github.com/JinPLu/WRBench