datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.kv-cache-compression-mbe
Matched-Budget Evaluation (MBE) — KV Cache Compression
A standardized reporting protocol for KV cache compression in LLM inference. MBE is
not a new task benchmark; it is a thin reporting layer that fixes which models, tasks,
and budgets results are reported at, so that numbers from different papers become
comparable.
Manifest (mbe_manifest.json): the frozen evaluation specification — model suite,
task suite (consuming existing benchmarks: LongBench, RULER, SCBench, GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/Rohithreddybc/kv-cache-compression-mbe.
