WRBench/wrbench-human-annotations
WRBench Human Annotations This dataset contains the human comparison labels used to validate WRBench's automatic evaluation metrics. Version Update: 2026-07-07 We updated the release after rechecking videos that changed during benchmark maintenance. The release now includes: 1,741 clean comparison rows. 4,302 individual human judgments. 585 newly rechecked current-benchmark comparisons, each reviewed by three annotators. Majority-label summaries for the newly… See the full description on the dataset page: https://huggingface.co/datasets/WRBench/wrbench-human-annotations.
WRBench Human Annotations
This dataset contains the human comparison labels used to validate WRBench's automatic evaluation metrics.
Version Update: 2026-07-07
We updated the release after rechecking videos that changed during benchmark maintenance. The release now includes:
- 1,741 clean comparison rows.
- 4,302 individual human judgments.
- 585 newly rechecked current-benchmark comparisons, each reviewed by three annotators.
- Majority-label summaries for the newly rechecked subset.
These counts describe this dated human-label release. Moving leaderboard, video, and model counts are reported by the results and videos datasets.
Videos for the current recheck
All 1,170 endpoints in current_review_consensus resolve to the exact bytes shown to annotators. Of these, 897 are in wrbench-videos; the remaining 273 endpoints refer to 106 unique files bundled here under videos/ and indexed by asset_manifest.parquet.
The bundled files are dated human-review evidence. They use the earlier profile-unsuffixed IDs and are not additional leaderboard videos or aliases for the 30-degree and 60-degree benchmark profiles. Join on video_asset_id, whose sha256: value verifies the file bytes.
Configs
pairs: clean comparison rows used by the public WRBench release.annotation_to_video: maps each comparison to the corresponding video assets.current_review_consensus: majority labels for the 585 newly rechecked comparisons.current_review_submissions: anonymized per-annotator labels for those 585 comparisons.current_review_statusandexclusions: records kept for transparency, but not used as clean comparison labels.review_assets: paths, byte counts, and SHA256 identifiers for the 106 review-only videos bundled with this dataset.
Annotator identities are anonymized in the public files.
Links
- Videos: https://huggingface.co/datasets/WRBench/wrbench-videos
- Results: https://huggingface.co/datasets/WRBench/wrbench-results
- Prompts: https://huggingface.co/datasets/WRBench/wrbench-natural25
- Leaderboard: https://huggingface.co/spaces/WRBench/wrbench-leaderboard
- Collection: https://huggingface.co/collections/WRBench/wrbench-current-world-models-lack-a-persistent-state-core
- Paper: https://arxiv.org/abs/2606.20545
- Code: https://github.com/JinPLu/WRBench
