CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zai-org /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongBench-v2.textmultiple-choicen<1K56 likes87k downloads2y agoHugging Face02recursal /longbench-v2 Citation @article{bai2024longbench2, title={LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks}, author={Yushi Bai and Shangqing Tu and Jiajie Zhang and Hao Peng and Xiaozhi Wang and Xin Lv and Shulin Cao and Jiazheng Xu and Lei Hou and Yuxiao Dong and Jie Tang and Juanzi Li}, journal={arXiv preprint arXiv:2412.15204}, year={2024} } multiple-choicen<1K0 likes2.4k downloads1y agoHugging Face03elbisreverri /longbenchv2_fc_labellevel_check0 likes233 downloads8mo agoHugging Face04MrBigBrane /LongBench-v2-32k-CoTtextn<1K0 likes170 downloads17d agoHugging Face05simonjegou /LongBench-v2textn<1K0 likes159 downloads10mo agoHugging Face06MrBigBrane /longbench-v2-32ktextn<1K0 likes105 downloads15d agoHugging Face07MrBigBrane /longbench-v2-shorttextn<1K0 likes70 downloads15d agoHugging Face08elbisreverri /longbenchv2_qa_directly0 likes60 downloads8mo agoHugging Face09itsnamgyu /longbench_synthetic_v2 LongBench Synthetic V2 tabular10K<n<100K0 likes60 downloads6mo agoHugging Face10leideng /longbench-v2-view LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-v2-view.textmultiple-choice1K<n<10K0 likes59 downloads6mo agoHugging Face11Green-Sky /LongBench-v2-for-llama.cppLongBench v2 converted for the llama.cpp perplexity multiple choice tool. [!WARNING] !! Currently does not work, will fix it in the near future. Probably. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems… See the full description on the dataset page: https://huggingface.co/datasets/Green-Sky/LongBench-v2-for-llama.cpp.textmultiple-choicen<1K0 likes49 downloads6mo agoHugging Face12lindsay21 /longbench_v2_transformed_rlThis dataset is introduced in arxiv.org/abs/2602.12108 BibTeX: @misc{liu2026pensieveparadigmstatefullanguage, title={The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context}, author={Xiaoyuan Liu and Tian Liang and Dongyang Ma and Deyu Zhou and Haitao Mi and Pinjia He and Yan Wang}, year={2026}, eprint={2602.12108}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/abs/2602.12108}, } textn<1K1 likes39 downloads8mo agoHugging Face13giulio98 /LongBench-v2text1K<n<10K0 likes33 downloads5mo agoHugging Face14HanyueShen /YunXiaoHe-LongBench-v2-Comparison YunXiaoHe on LongBench-v2 YunXiaoHe answered 337 of 503 LongBench-v2 questions correctly, or 67.0%. Four unfinished questions count as incorrect. This is a provisional, self-reported result rather than an official leaderboard submission. The left panel puts that result beside the LongBench-v2 site's first nine model rows, ranked by its overall chain-of-thought score, and all six LongBench-v2 configurations in the 2026 Prime Agent study. The official-site rows range from 56.0%… See the full description on the dataset page: https://huggingface.co/datasets/HanyueShen/YunXiaoHe-LongBench-v2-Comparison.tabularn<1K0 likes25 downloads2d agoHugging Face15giulio98 /LongBench-v2-1024textn<1K0 likes23 downloads1y agoHugging Face16JamesBegin /LongBench-v2-Pause1 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/JamesBegin/LongBench-v2-Pause1.textmultiple-choicen<1K1 likes18 downloads2y agoHugging Face17Xnhyacinth /LongBench-v2text1K<n<10K0 likes18 downloads2y agoHugging Face18xiaoyuanliu /LongBench-v2-verifiedtextn<1K0 likes16 downloads8mo agoHugging Face19zwhe99 /LongBench-v2-reformattedtextn<1K0 likes15 downloads1y agoHugging Face20giulio98 /LongBench-v2-newtextn<1K0 likes15 downloads5mo agoHugging Face21xiaoyuanliu /LongBench-v2-rlvrtextn<1K0 likes14 downloads9mo agoHugging Face22AlioLeuchtmann /LongBench-v2-with-len-in-tokenIncluding Length in Tokens: { "Qwen/Qwen3-0.6B": 283999 } Currently only for Qwen3 Models Citation @article{bai2024longbench2, title={LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks}, author={Yushi Bai and Shangqing Tu and Jiajie Zhang and Hao Peng and Xiaozhi Wang and Xin Lv and Shulin Cao and Jiazheng Xu and Lei Hou and Yuxiao Dong and Jie Tang and Juanzi Li}, journal={arXiv preprint arXiv:2412.15204}, year={2024} } textn<1K0 likes12 downloads8mo agoHugging Face23nizarislah /longbenchv2-topk-qwen7b-fixedtabularn<1K0 likes11 downloads1y agoHugging Face24xiaoyuanliu /LongBench-v2-smalltextn<1K0 likes10 downloads1y agoHugging Face25lillycyx /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/lillycyx/LongBench-v2.textmultiple-choicen<1K0 likes10 downloads5mo agoHugging Face26MinX125 /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/MinX125/LongBench-v2.textmultiple-choicen<1K0 likes8 downloads7mo agoHugging Face27giulio98 /LongBench-v2-2048textn<1K0 likes7 downloads1y agoHugging Face28xiaoyuanliu /LongBench-v2-T100textn<1K0 likes7 downloads9mo agoHugging Face29giulio98 /LongBench-v2-8192textn<1K0 likes5 downloads1y agoHugging Face30giulio98 /LongBench-v2-16384textn<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.