CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zai-org /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongBench-v2.textmultiple-choicen<1K56 likes86k downloads2y agoHugging Face02caskcsg /LongBench-Pro LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark          LongBench-Pro, containing 1,500 samples, is entirely built on authentic, natural long documents and includes 11 primary tasks and 25 secondary tasks, covering all long-context capabilities assessed by existing benchmarks. It employs diverse evaluation metrics, enabling a more fine-grained measurement of model abilities, and provides a balanced set of bilingual samples in both… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/LongBench-Pro.textquestion-answering1K<n<10K9 likes2.5k downloads9mo agoHugging Face03HanyueShen /YunXiaoHe-LongBench-Eval 云小鹤 0.3.7 LongBench zero-shot 评测 This repository contains the public evidence for a complete 200-example LongBench v1 HotpotQA run by 云小鹤 (YunXiaoHe) 0.3.7. The release covers the evaluation result, item-level trace, usage records, figures and recomputation code. The proprietary agent implementation is outside the release. 云小鹤以 zero-shot 方式完成了全部 200 题。运行前没有针对 HotpotQA 进行专项训练、微调、示例拟合、阈值搜索或评测集优化。 Join the open technical review to inspect the scoring protocol, propose an… See the full description on the dataset page: https://huggingface.co/datasets/HanyueShen/YunXiaoHe-LongBench-Eval.documentquestion-answeringn<1K1 likes205 downloads24d agoHugging Face04leideng /longbench-view Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.textquestion-answering1K<n<10K0 likes194 downloads6mo agoHugging Face05fang0608 /LongBench-Pro LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark          LongBench-Pro, containing 1,500 samples, is entirely built on authentic, natural long documents and includes 11 primary tasks and 25 secondary tasks, covering all long-context capabilities assessed by existing benchmarks. It employs diverse evaluation metrics, enabling a more fine-grained measurement of model abilities, and provides a balanced set of bilingual samples in both… See the full description on the dataset page: https://huggingface.co/datasets/fang0608/LongBench-Pro.textquestion-answering1K<n<10K0 likes62 downloads5mo agoHugging Face06virgilR /LongBench-Pro LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark          LongBench-Pro, containing 1,500 samples, is entirely built on authentic, natural long documents and includes 11 primary tasks and 25 secondary tasks, covering all long-context capabilities assessed by existing benchmarks. It employs diverse evaluation metrics, enabling a more fine-grained measurement of model abilities, and provides a balanced set of bilingual samples in both… See the full description on the dataset page: https://huggingface.co/datasets/virgilR/LongBench-Pro.textquestion-answering1K<n<10K0 likes55 downloads5mo agoHugging Face07Green-Sky /LongBench-v2-for-llama.cppLongBench v2 converted for the llama.cpp perplexity multiple choice tool. [!WARNING] !! Currently does not work, will fix it in the near future. Probably. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems… See the full description on the dataset page: https://huggingface.co/datasets/Green-Sky/LongBench-v2-for-llama.cpp.textmultiple-choicen<1K0 likes47 downloads6mo agoHugging Face08JamesBegin /LongBench-v2-Pause1 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/JamesBegin/LongBench-v2-Pause1.textmultiple-choicen<1K1 likes19 downloads2y agoHugging Face09lillycyx /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/lillycyx/LongBench-v2.textmultiple-choicen<1K0 likes8 downloads5mo agoHugging Face10MinX125 /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/MinX125/LongBench-v2.textmultiple-choicen<1K0 likes7 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.