datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LongBenchLongBenchlongbench-longctx
longbench-longctx
Long-context requests for end-to-end LLM inference benchmarking in fastkernels — Scenario B. Exercises the regimes the bulk set can't reach: long-sequence attention (incl. sparse / sliding-window / DSA), RoPE/YaRN scaling, and large-KV decode.
What it's for
64 real long documents truncated into clean prefill-length buckets from 8K to 128K, each paired with its real multiple-choice question. Prefill-dominated: it measures how kernels scale with… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/longbench-longctx.YunXiaoHe-LongBench-Eval
云小鹤 0.3.7 LongBench zero-shot 评测
This repository contains the public evidence for a complete 200-example LongBench v1 HotpotQA run by 云小鹤 (YunXiaoHe) 0.3.7. The release covers the evaluation result, item-level trace, usage records, figures and recomputation code. The proprietary agent implementation is outside the release.
云小鹤以 zero-shot 方式完成了全部 200 题。运行前没有针对 HotpotQA 进行专项训练、微调、示例拟合、阈值搜索或评测集优化。
Join the open technical review to inspect the scoring protocol, propose an… See the full description on the dataset page: https://huggingface.co/datasets/HanyueShen/YunXiaoHe-LongBench-Eval.longbench_synthetic_v3_1
LongBench Synthetic V3.1
Dataset statistics
Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B.
Pool
Subset
Unique ctx
Sample rows
<8K
8-16K
16-32K
>32K
Median tok
p90 tok
Max tok
eval
hotpotqa
200
200
27
110
63
0
14,982
16,947
17,578
eval
hotpotqa_e
286
286
116
139
31
0
9,575
16,434
17,322
eval
musique
200
200
3
46
151
0
16,733
17… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v3_1.longbench_synthetic_v4_1
LongBench Synthetic V4.1
Dataset statistics
v4.1 = lbs_v4 train+val (verbatim) + a test split absorbing every config from dac-research/extra_evals_v1 not already in v4. Five ZeroScrolls configs whose upstream corpora collide with lbs_v3.1 train/test (gov_report, qmsum, qasper, narrative_qa, musique) were dropped outright. All test rows have rubrics backfilled via the longbench_generate_rubric_for_imported.jinja template (same path as lbs_v3.1 / lbs_v4 Stage 1b).… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v4_1.LongBench-newLongBench
LongBench
Dataset Summary
LongBench is a bilingual, multitask benchmark for evaluating long-context understanding in large language models. It covers long-text application scenarios including single-document question answering, multi-document question answering, summarization, few-shot learning, synthetic long-context tasks, and code completion.
This Hugging Face dataset repository repackages locally downloaded LongBench JSONL files into a clean, typed, data-only… See the full description on the dataset page: https://huggingface.co/datasets/GinkgoQ/LongBench.ptb-longbenchwrite100-LongBenchLongBench-elongbench_synthetic_v2
LongBench Synthetic V2
longbench_kvpressLongBench-512LongBench-2048LongBench-sLongBench-BM25-512LongBenchLongBench-1024LongBench-BM25-1024LongBench-BM25-2048longbench_synthetic_v1
LongBench Synthetic V1
longbench_synthetic_v3
LongBench Synthetic V3
Dataset statistics
Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B.
Pool
Subset
Unique ctx
Sample rows
<8K
8-16K
16-32K
>32K
Median tok
p90 tok
Max tok
eval
2wikimqa
197
197
139
54
4
0
6,627
13,271
16,982
eval
gov_report200
200
89
84
24
3
8,902
18,212
52,521
eval
musique
200
200
3
46
151
0
16,733
17,231
17,824… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v3.LongBench-4096rlm-longbenchpro-rawLongBench-lccLongBench-Qwenlongbench_synthetic_v4
LongBench Synthetic V4
Dataset statistics
Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B.
Subset
Unique ctx
Sample rows
<8K
8-16K
16-32K
32-64K
64-128K
>128K
Median tok
p90 tok
Max tok
longbench_v2_multidoc_qa
123
984
0
104
152
264
152
312
61,425270,817
960,383
longbench_v2_singledoc_qa
166
1,400
0
112
376
128
352
432
75,354
221,060
865… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v4.longbenchv2-topk-qwen7b-fixedLongBench-v2
