CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zai-org /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongBench-v2.textmultiple-choicen<1K56 likes86k downloads2y agoHugging Face02zai-org /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.question-answering1K<n<10K191 likes57k downloads2y agoHugging Face03caskcsg /LongBench-Pro LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark          LongBench-Pro, containing 1,500 samples, is entirely built on authentic, natural long documents and includes 11 primary tasks and 25 secondary tasks, covering all long-context capabilities assessed by existing benchmarks. It employs diverse evaluation metrics, enabling a more fine-grained measurement of model abilities, and provides a balanced set of bilingual samples in both… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/LongBench-Pro.textquestion-answering1K<n<10K9 likes2.5k downloads9mo agoHugging Face04recursal /longbench-v2 Citation @article{bai2024longbench2, title={LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks}, author={Yushi Bai and Shangqing Tu and Jiajie Zhang and Hao Peng and Xiaozhi Wang and Xin Lv and Shulin Cao and Jiazheng Xu and Lei Hou and Yuxiao Dong and Jie Tang and Juanzi Li}, journal={arXiv preprint arXiv:2412.15204}, year={2024} } multiple-choicen<1K0 likes2.4k downloads1y agoHugging Face05jannalu /LongBench Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/jannalu/LongBench.textquestion-answering1K<n<10K0 likes235 downloads11mo agoHugging Face06llm-jp /llm-jp-longbench-JEMHop llm-jp-longbench-JEMHopQA llm-jp LongBench ベンチマークについて このデータセットは,GitHub リポジトリhttps://github.com/llm-jp/llm-jp-longbenchで公開されているllm-jp LongBenchベンチマークの評価対象データセットの一部として構築されています。 llm-jp LongBench ベンチマークは,日本語大型言語モデル(LLM)のロングコンテキスト処理能力を体系的に評価することを目的としており,複数の長文コンテキスト QA データセットを含んでいます。 本データセットはその一つです。 データセット概要 本データセットは、日本語の説明可能マルチホップ質問応答データセットJEMHopQA (Ishii et al., 2024)を基に、Wikipedia記事を付与することで構築したロングコンテキストQA評価用データセットです。 最大65… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-longbench-JEMHop.textquestion-answeringn<1K1 likes216 downloads7mo agoHugging Face07llm-jp /llm-jp-longbench-NIILC llm-jp-longbench-NIILC llm-jp LongBench ベンチマークについて このデータセットは,GitHub リポジトリhttps://github.com/llm-jp/llm-jp-longbenchで公開されているllm-jp LongBenchベンチマークの評価対象データセットの一部として構築されています。 llm-jp LongBench ベンチマークは,日本語大型言語モデル(LLM)のロングコンテキスト処理能力を体系的に評価することを目的としており,複数の長文コンテキスト QA データセットを含んでいます。 本データセットはその一つです。 データセット概要 本データセットは,日本語質問応答データセット NIILC (Sekine, 2003)を基に, 回答が一意に定まり,かつ時間によって正解が変化しない質問のみを選別し, それらに対応する Wikipedia 記事をコンテキストとして付与することで構築した, ロングコンテキスト QA 評価用データセットです。… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-longbench-NIILC.textquestion-answeringn<1K1 likes207 downloads7mo agoHugging Face08HanyueShen /YunXiaoHe-LongBench-Eval 云小鹤 0.3.7 LongBench zero-shot 评测 This repository contains the public evidence for a complete 200-example LongBench v1 HotpotQA run by 云小鹤 (YunXiaoHe) 0.3.7. The release covers the evaluation result, item-level trace, usage records, figures and recomputation code. The proprietary agent implementation is outside the release. 云小鹤以 zero-shot 方式完成了全部 200 题。运行前没有针对 HotpotQA 进行专项训练、微调、示例拟合、阈值搜索或评测集优化。 Join the open technical review to inspect the scoring protocol, propose an… See the full description on the dataset page: https://huggingface.co/datasets/HanyueShen/YunXiaoHe-LongBench-Eval.documentquestion-answeringn<1K1 likes205 downloads24d agoHugging Face09leideng /longbench-view Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.textquestion-answering1K<n<10K0 likes194 downloads6mo agoHugging Face10bzantium /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.textquestion-answering1K<n<10K1 likes187 downloads3y agoHugging Face11yanbingzheng /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.question-answering1K<n<10K2 likes133 downloads3y agoHugging Face12GinkgoQ /LongBench LongBench Dataset Summary LongBench is a bilingual, multitask benchmark for evaluating long-context understanding in large language models. It covers long-text application scenarios including single-document question answering, multi-document question answering, summarization, few-shot learning, synthetic long-context tasks, and code completion. This Hugging Face dataset repository repackages locally downloaded LongBench JSONL files into a clean, typed, data-only… See the full description on the dataset page: https://huggingface.co/datasets/GinkgoQ/LongBench.tabularquestion-answering1K<n<10K1 likes120 downloads4mo agoHugging Face13hyg444 /LongBench Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/hyg444/LongBench.question-answering1K<n<10K0 likes62 downloads5mo agoHugging Face14fang0608 /LongBench-Pro LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark          LongBench-Pro, containing 1,500 samples, is entirely built on authentic, natural long documents and includes 11 primary tasks and 25 secondary tasks, covering all long-context capabilities assessed by existing benchmarks. It employs diverse evaluation metrics, enabling a more fine-grained measurement of model abilities, and provides a balanced set of bilingual samples in both… See the full description on the dataset page: https://huggingface.co/datasets/fang0608/LongBench-Pro.textquestion-answering1K<n<10K0 likes62 downloads5mo agoHugging Face15leideng /longbench-v2-view LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-v2-view.textmultiple-choice1K<n<10K0 likes58 downloads6mo agoHugging Face16virgilR /LongBench-Pro LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark          LongBench-Pro, containing 1,500 samples, is entirely built on authentic, natural long documents and includes 11 primary tasks and 25 secondary tasks, covering all long-context capabilities assessed by existing benchmarks. It employs diverse evaluation metrics, enabling a more fine-grained measurement of model abilities, and provides a balanced set of bilingual samples in both… See the full description on the dataset page: https://huggingface.co/datasets/virgilR/LongBench-Pro.textquestion-answering1K<n<10K0 likes55 downloads5mo agoHugging Face17budecosystem /longbench_hotpotqa Bud Ecosystem mirror of THUDM/LongBench — a verbatim copy for offline, reproducible model evaluation. License unchanged (CC-BY-SA-4.0 (inherited from HotpotQA source; LongBench code is MIT)); all rights remain with the original authors. Share-alike: any derivative must stay under the same CC-BY-SA license. Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models.… See the full description on the dataset page: https://huggingface.co/datasets/budecosystem/longbench_hotpotqa.question-answering1K<n<10K0 likes50 downloads2mo agoHugging Face18Green-Sky /LongBench-v2-for-llama.cppLongBench v2 converted for the llama.cpp perplexity multiple choice tool. [!WARNING] !! Currently does not work, will fix it in the near future. Probably. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems… See the full description on the dataset page: https://huggingface.co/datasets/Green-Sky/LongBench-v2-for-llama.cpp.textmultiple-choicen<1K0 likes47 downloads6mo agoHugging Face19minliii /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.question-answering1K<n<10K0 likes23 downloads9mo agoHugging Face20JamesBegin /LongBench-v2-Pause1 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/JamesBegin/LongBench-v2-Pause1.textmultiple-choicen<1K1 likes19 downloads2y agoHugging Face21lillycyx /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/lillycyx/LongBench-v2.textmultiple-choicen<1K0 likes8 downloads5mo agoHugging Face22MinX125 /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/MinX125/LongBench-v2.textmultiple-choicen<1K0 likes7 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.