datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
example_10kbp_human_annotationsexample_eval_only_10kb_human_annotationswrbench-human-annotations
WRBench Human Annotations
This dataset contains the human comparison labels used to validate WRBench's
automatic evaluation metrics.
Version Update: 2026-07-07
We updated the release after rechecking videos that changed during benchmark
maintenance. The release now includes:
1,741 clean comparison rows.
4,302 individual human judgments.
585 newly rechecked current-benchmark comparisons, each reviewed by three
annotators.
Majority-label summaries for the newly… See the full description on the dataset page: https://huggingface.co/datasets/WRBench/wrbench-human-annotations.aya23-human-annotationsrobust_long_abstractive_human_annotationOriginal repository
How Far are We from Robust Long Abstractive Summarization? (EMNLP 2022)
[Paper]
Huan Yee Koh*, Jiaxin Ju*, He Zhang, Ming Liu, Shirui Pan
(* denotes equal contribution)
Human Annotation of Model-Generated Summaries
Data Field
Definition
dataset
Whether the model-generated summary is from arXiv or GovReport dataset.
dataset_id
ID_ + document ID of the dataset. To match the IDs with original datasets, please remove the "ID_"… See the full description on the dataset page: https://huggingface.co/datasets/gigant/robust_long_abstractive_human_annotation.
