datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.stillwarm-kv-cache-artifact
A downloadable KV-cache save file — with the honest math
One llama-server slot save: the first 8,192 Llama-tokens of Frankenstein
(public domain), prefilled by Qwen2.5-7B-Instruct Q4_K_M (Apache-2.0 model —
chosen over Llama specifically for artifact licensing) and saved with a
stillwarm sidecar.
This file is USELESS unless your setup matches the sidecar exactly:
field
value
llama.cpp build
b9871 (ef2d770117db45b05aa7ecd1b0acca36370c5470) — advisory: ±5 weeks measured… See the full description on the dataset page: https://huggingface.co/datasets/vimalnakrani/stillwarm-kv-cache-artifact.kvcache-bench-results
Mingxin KV-Cache Tiered-Storage Benchmark Results
Measured results for LLM KV Cache tiered-storage acceleration on 8× AMD Instinct MI308X (ROCm 7.2, vLLM 0.20.1+rocm721, LMCache upstream mainline), model Qwen3-Coder-480B-FP8, published by Mingxin Technology.
Device under test: Mingxin FX100 all-flash NVMe-oF array (4-disk RAID0, RoCEv2, 100 GbE) vs local NVMe (PCIe Gen4) vs recompute-only baseline.
Files
kvcache_bench_results.json — all experiments in structured… See the full description on the dataset page: https://huggingface.co/datasets/wangqiyuan2026/kvcache-bench-results.kv_cache_logits_exp2kv-cache-compression-mbe
Matched-Budget Evaluation (MBE) — KV Cache Compression
A standardized reporting protocol for KV cache compression in LLM inference. MBE is
not a new task benchmark; it is a thin reporting layer that fixes which models, tasks,
and budgets results are reported at, so that numbers from different papers become
comparable.
Manifest (mbe_manifest.json): the frozen evaluation specification — model suite,
task suite (consuming existing benchmarks: LongBench, RULER, SCBench, GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/Rohithreddybc/kv-cache-compression-mbe.
