CoolFace
16 results

longbench

zai-org /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongBench-v2.textmultiple-choicen<1K56 likes86k downloads2y agoHugging Facezai-org /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.question-answering1K<n<10K191 likes58k downloads2y agoHugging FaceXnhyacinth /LongBenchtabular1K<n<10K7 likes33k downloads1y agoHugging Facecaskcsg /LongBench-Pro LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark          LongBench-Pro, containing 1,500 samples, is entirely built on authentic, natural long documents and includes 11 primary tasks and 25 secondary tasks, covering all long-context capabilities assessed by existing benchmarks. It employs diverse evaluation metrics, enabling a more fine-grained measurement of model abilities, and provides a balanced set of bilingual samples in both… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/LongBench-Pro.textquestion-answering1K<n<10K9 likes2.5k downloads9mo agoHugging Facerecursal /longbench-v2 Citation @article{bai2024longbench2, title={LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks}, author={Yushi Bai and Shangqing Tu and Jiajie Zhang and Hao Peng and Xiaozhi Wang and Xin Lv and Shulin Cao and Jiazheng Xu and Lei Hou and Yuxiao Dong and Jie Tang and Juanzi Li}, journal={arXiv preprint arXiv:2412.15204}, year={2024} } multiple-choicen<1K0 likes2.5k downloads1y agoHugging Facegiulio98 /LongBenchtabular1K<n<10K0 likes1.1k downloads1y agoHugging Face