datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LongBenchLongBenchlongbenchlongbench-longctx
longbench-longctx
Long-context requests for end-to-end LLM inference benchmarking in fastkernels — Scenario B. Exercises the regimes the bulk set can't reach: long-sequence attention (incl. sparse / sliding-window / DSA), RoPE/YaRN scaling, and large-KV decode.
What it's for
64 real long documents truncated into clean prefill-length buckets from 8K to 128K, each paired with its real multiple-choice question. Prefill-dominated: it measures how kernels scale with… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/longbench-longctx.LongBench
Introduction
LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/jannalu/LongBench.LongBench-v2-32k-CoTLongBench-v2longbench_synthetic_v3_1
LongBench Synthetic V3.1
Dataset statistics
Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B.
Pool
Subset
Unique ctx
Sample rows
<8K
8-16K
16-32K
>32K
Median tok
p90 tok
Max tok
eval
hotpotqa
200
200
27
110
63
0
14,982
16,947
17,578
eval
hotpotqa_e
286
286
116
139
31
0
9,575
16,434
17,322
eval
musique
200
200
3
46
151
0
16,733
17… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v3_1.longbench_synthetic_v4_1
LongBench Synthetic V4.1
Dataset statistics
v4.1 = lbs_v4 train+val (verbatim) + a test split absorbing every config from dac-research/extra_evals_v1 not already in v4. Five ZeroScrolls configs whose upstream corpora collide with lbs_v3.1 train/test (gov_report, qmsum, qasper, narrative_qa, musique) were dropped outright. All test rows have rubrics backfilled via the longbench_generate_rubric_for_imported.jinja template (same path as lbs_v3.1 / lbs_v4 Stage 1b).… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v4_1.LongBench-newLongBenchLongBench-multilingualWIP, please don't use yet
LongBench
LongBench
Dataset Summary
LongBench is a bilingual, multitask benchmark for evaluating long-context understanding in large language models. It covers long-text application scenarios including single-document question answering, multi-document question answering, summarization, few-shot learning, synthetic long-context tasks, and code completion.
This Hugging Face dataset repository repackages locally downloaded LongBench JSONL files into a clean, typed, data-only… See the full description on the dataset page: https://huggingface.co/datasets/GinkgoQ/LongBench.longbench-v2-32kptb-longbenchwritelongbench-v2-short100-LongBenchLongBench-elongbenchlongbench-narrativeqalongbench_synthetic_v2
LongBench Synthetic V2
longbench_kvpressLongBench_bgLongBench-512LongBench-2048LongBench-sLongBench-v2longbench_v2_transformed_rlThis dataset is introduced in arxiv.org/abs/2602.12108
BibTeX:
@misc{liu2026pensieveparadigmstatefullanguage,
title={The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context},
author={Xiaoyuan Liu and Tian Liang and Dongyang Ma and Deyu Zhou and Haitao Mi and Pinjia He and Yan Wang},
year={2026},
eprint={2602.12108},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2602.12108},
}
olmoe3_stage3_longbench_evalLongBench-v2-1024
