datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UPBench-Error-verified-v2
UPBench-Error-verified-v2
LCZZZZ/UPBench-Error → generation_error 子集,经两轮人工核验后保留的
1,219 条样本。每条样本视觉上看不出明显的低级生成缺陷。
筛选过程
步骤
剩余
原始 generation_error 样本
5,761
剔除 is_gui=true(GUI-World / egoproactive 屏幕录制)
3,390
第一轮:逐条过目 error_clip.mp4
good 1,420 / bad 1,970
第二轮:对第一轮 good 再过一遍
good 1,219 / bad 201
第二轮刷掉了第一轮 14.2% 的样本,最终保留率 1,219 / 3,390 = 36.0%。
内容
metadata/manifest-verified.jsonl 1,219 条,原 manifest 全部 25 个字段逐字保留,… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/UPBench-Error-verified-v2.UPBench-Error-verified
UPBench-Error-verified
人工核验过的 LCZZZZ/UPBench-Error → generation_error 子集。
保留其中看不出明显低级生成缺陷的样本。
这是怎么筛出来的
从原数据集 5,761 条 generation_error 样本出发:
步骤
剩余
原始样本
5,761
剔除 is_gui=true(GUI-World / egoproactive 屏幕录制)
3,390
逐条人工过目 error_clip.mp4
3,390 全部看完
判定为无明显生成缺陷(good)
1,420
判定为有缺陷(bad)
1,970
内容
metadata/manifest-verified.jsonl 1,420 条,原 manifest 字段逐字保留,
另加 human_verification 字段… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/UPBench-Error-verified.
