CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zai-org /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.question-answering1K<n<10K191 likes58k downloads2y agoHugging Face02jannalu /LongBench Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/jannalu/LongBench.textquestion-answering1K<n<10K0 likes232 downloads11mo agoHugging Face03sfc-gh-goliaro /longbench-longctx longbench-longctx Long-context requests for end-to-end LLM inference benchmarking in fastkernels — Scenario B. Exercises the regimes the bulk set can't reach: long-sequence attention (incl. sparse / sliding-window / DSA), RoPE/YaRN scaling, and large-KV decode. What it's for 64 real long documents truncated into clean prefill-length buckets from 8K to 128K, each paired with its real multiple-choice question. Prefill-dominated: it measures how kernels scale with… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/longbench-longctx.tabulartext-generationn<1K0 likes208 downloads3mo agoHugging Face04leideng /longbench-view Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.textquestion-answering1K<n<10K0 likes205 downloads6mo agoHugging Face05bzantium /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.textquestion-answering1K<n<10K1 likes185 downloads3y agoHugging Face06yanbingzheng /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.question-answering1K<n<10K2 likes132 downloads3y agoHugging Face07GinkgoQ /LongBench LongBench Dataset Summary LongBench is a bilingual, multitask benchmark for evaluating long-context understanding in large language models. It covers long-text application scenarios including single-document question answering, multi-document question answering, summarization, few-shot learning, synthetic long-context tasks, and code completion. This Hugging Face dataset repository repackages locally downloaded LongBench JSONL files into a clean, typed, data-only… See the full description on the dataset page: https://huggingface.co/datasets/GinkgoQ/LongBench.tabularquestion-answering1K<n<10K1 likes120 downloads4mo agoHugging Face08hyg444 /LongBench Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/hyg444/LongBench.question-answering1K<n<10K0 likes61 downloads5mo agoHugging Face09budecosystem /longbench_hotpotqa Bud Ecosystem mirror of THUDM/LongBench — a verbatim copy for offline, reproducible model evaluation. License unchanged (CC-BY-SA-4.0 (inherited from HotpotQA source; LongBench code is MIT)); all rights remain with the original authors. Share-alike: any derivative must stay under the same CC-BY-SA license. Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models.… See the full description on the dataset page: https://huggingface.co/datasets/budecosystem/longbench_hotpotqa.question-answering1K<n<10K0 likes50 downloads2mo agoHugging Face10Syon-Li /LongbenchSeg Dataset Card for LongbenchSeg The segmented version of Longbench using the trained segmenter from the paper Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation. We set the recursion depth to 1 and use the threshold value of 0.4. Dataset Details Dataset Description The newly introduced segmentation columns are: chunks: The segmented chunks. cut_prob: The corresponding segmenting probability for each… See the full description on the dataset page: https://huggingface.co/datasets/Syon-Li/LongbenchSeg.text-generation1K<n<10K0 likes34 downloads8d agoHugging Face11minliii /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.question-answering1K<n<10K0 likes22 downloads9mo agoHugging Face12viktor-shcherb /longbench2-128k-plus LongBench2-128k-plus LongBench2-128k-plus is a long-context corpus derived from the zai-org/LongBench-v2 benchmark. It keeps only the "long" examples and exposes just the raw long documents, making it convenient for: long-context pretraining or continued training, long-context adaptation (e.g., RoPE scaling, attention tuning), retrieval and RAG-style experimentation where only documents are needed. All question/answer and multiple-choice metadata from LongBench v2 are dropped;… See the full description on the dataset page: https://huggingface.co/datasets/viktor-shcherb/longbench2-128k-plus.texttext-generationn<1K0 likes15 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.