microsoft/XL-DocBench
XL-DocBench Evidence-grounded reasoning across hundreds or thousands of pages. Fully verified by 194 human experts. Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2 1Wuhan University 2Microsoft †Equal contribution ‡Work done during an internship at MSRA *Project leader Project Page · Paper · Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.
7932
1{"question_id":"adubench_single_000001","prediction":"the biggest single risk to human health worldwide"}2{"question_id":"adubench_single_000002","prediction":"25"}3{"question_id":"adubench_single_000083","prediction":"B"}4{"question_id":"adubench_single_000193","prediction":"Not answerable"}5{"question_id":"adubench_cross_000001","prediction":"macroprudential measures"}6 