CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Xnhyacinth /LongBenchtabular1K<n<10K7 likes32k downloads1y agoHugging Face02giulio98 /LongBenchtabular1K<n<10K0 likes1.1k downloads1y agoHugging Face03sfc-gh-goliaro /longbench-longctx longbench-longctx Long-context requests for end-to-end LLM inference benchmarking in fastkernels — Scenario B. Exercises the regimes the bulk set can't reach: long-sequence attention (incl. sparse / sliding-window / DSA), RoPE/YaRN scaling, and large-KV decode. What it's for 64 real long documents truncated into clean prefill-length buckets from 8K to 128K, each paired with its real multiple-choice question. Prefill-dominated: it measures how kernels scale with… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/longbench-longctx.tabulartext-generationn<1K0 likes260 downloads3mo agoHugging Face04HanyueShen /YunXiaoHe-LongBench-Eval 云小鹤 0.3.7 LongBench zero-shot 评测 This repository contains the public evidence for a complete 200-example LongBench v1 HotpotQA run by 云小鹤 (YunXiaoHe) 0.3.7. The release covers the evaluation result, item-level trace, usage records, figures and recomputation code. The proprietary agent implementation is outside the release. 云小鹤以 zero-shot 方式完成了全部 200 题。运行前没有针对 HotpotQA 进行专项训练、微调、示例拟合、阈值搜索或评测集优化。 Join the open technical review to inspect the scoring protocol, propose an… See the full description on the dataset page: https://huggingface.co/datasets/HanyueShen/YunXiaoHe-LongBench-Eval.documentquestion-answeringn<1K1 likes205 downloads24d agoHugging Face05dac-research /longbench_synthetic_v3_1 LongBench Synthetic V3.1 Dataset statistics Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B. Pool Subset Unique ctx Sample rows <8K 8-16K 16-32K >32K Median tok p90 tok Max tok eval hotpotqa 200 200 27 110 63 0 14,982 16,947 17,578 eval hotpotqa_e 286 286 116 139 31 0 9,575 16,434 17,322 eval musique 200 200 3 46 151 0 16,733 17… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v3_1.tabular10K<n<100K1 likes161 downloads4mo agoHugging Face06dac-research /longbench_synthetic_v4_1 LongBench Synthetic V4.1 Dataset statistics v4.1 = lbs_v4 train+val (verbatim) + a test split absorbing every config from dac-research/extra_evals_v1 not already in v4. Five ZeroScrolls configs whose upstream corpora collide with lbs_v3.1 train/test (gov_report, qmsum, qasper, narrative_qa, musique) were dropped outright. All test rows have rubrics backfilled via the longbench_generate_rubric_for_imported.jinja template (same path as lbs_v3.1 / lbs_v4 Stage 1b).… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v4_1.tabular10K<n<100K1 likes158 downloads4mo agoHugging Face07giulio98 /LongBench-newtabular1K<n<10K0 likes154 downloads9mo agoHugging Face08GinkgoQ /LongBench LongBench Dataset Summary LongBench is a bilingual, multitask benchmark for evaluating long-context understanding in large language models. It covers long-text application scenarios including single-document question answering, multi-document question answering, summarization, few-shot learning, synthetic long-context tasks, and code completion. This Hugging Face dataset repository repackages locally downloaded LongBench JSONL files into a clean, typed, data-only… See the full description on the dataset page: https://huggingface.co/datasets/GinkgoQ/LongBench.tabularquestion-answering1K<n<10K1 likes120 downloads4mo agoHugging Face09HaimingW /ptb-longbenchwritetabularn<1K0 likes77 downloads3mo agoHugging Face10albertgong1 /100-LongBenchtabular1K<n<10K0 likes69 downloads11mo agoHugging Face11Xnhyacinth /LongBench-etabular1K<n<10K1 likes68 downloads2y agoHugging Face12itsnamgyu /longbench_synthetic_v2 LongBench Synthetic V2 tabular10K<n<100K0 likes49 downloads6mo agoHugging Face13yuanfengustc /longbench_kvpresstabular1K<n<10K0 likes45 downloads2y agoHugging Face14giulio98 /LongBench-512tabular1K<n<10K0 likes42 downloads1y agoHugging Face15giulio98 /LongBench-2048tabular1K<n<10K0 likes42 downloads1y agoHugging Face16giulio98 /LongBench-stabular1K<n<10K0 likes41 downloads1y agoHugging Face17giulio98 /LongBench-BM25-512tabular1K<n<10K0 likes31 downloads1y agoHugging Face18jungi /LongBenchtabular1K<n<10K0 likes29 downloads1y agoHugging Face19giulio98 /LongBench-1024tabular1K<n<10K0 likes26 downloads1y agoHugging Face20giulio98 /LongBench-BM25-1024tabular1K<n<10K0 likes26 downloads1y agoHugging Face21giulio98 /LongBench-BM25-2048tabular1K<n<10K0 likes25 downloads1y agoHugging Face22itsnamgyu /longbench_synthetic_v1 LongBench Synthetic V1 tabular1K<n<10K0 likes24 downloads6mo agoHugging Face23dac-research /longbench_synthetic_v3 LongBench Synthetic V3 Dataset statistics Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B. Pool Subset Unique ctx Sample rows <8K 8-16K 16-32K >32K Median tok p90 tok Max tok eval 2wikimqa 197 197 139 54 4 0 6,627 13,271 16,982 eval gov_report200 200 89 84 24 3 8,902 18,212 52,521 eval musique 200 200 3 46 151 0 16,733 17,231 17,824… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v3.tabular10K<n<100K0 likes21 downloads5mo agoHugging Face24giulio98 /LongBench-4096tabular1K<n<10K0 likes20 downloads1y agoHugging Face25ShayanShamsi /rlm-longbenchpro-rawtabularn<1K0 likes15 downloads8mo agoHugging Face26giulio98 /LongBench-lcctabularn<1K0 likes14 downloads1y agoHugging Face27giulio98 /LongBench-Qwentabular1K<n<10K0 likes13 downloads1y agoHugging Face28dac-research /longbench_synthetic_v4 LongBench Synthetic V4 Dataset statistics Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B. Subset Unique ctx Sample rows <8K 8-16K 16-32K 32-64K 64-128K >128K Median tok p90 tok Max tok longbench_v2_multidoc_qa 123 984 0 104 152 264 152 312 61,425270,817 960,383 longbench_v2_singledoc_qa 166 1,400 0 112 376 128 352 432 75,354 221,060 865… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v4.tabular1K<n<10K0 likes13 downloads5mo agoHugging Face29nizarislah /longbenchv2-topk-qwen7b-fixedtabularn<1K0 likes11 downloads1y agoHugging Face30zypho /LongBench-v2tabularn<1K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.