kv_cache
TinyLlama-1.1B-Chat-v1.0-kvcache-fp8-tensorTinyLlama-1.1B-Chat-v1.0-kvcache-fp8-attn_headkv_cache_fp8-e2ekv_cache_gptq_tinyllama-e2eTinyLlama-1.1B-compressed-tensors-kv-cache-schemeopt-125m-gqa-ub-6-best-for-KV-cachefacebook-opt-6.7b-gqa-ub-16-best-for-KV-cachefacebook-opt-125m-qcqa-ub-6-best-for-KV-cache
Datasets
All datasets matching “kv_cache”daily-paper-2026-07-28-kv-cache-tiering-pd-disagg-cost
KV Cache Tiering Meets Prefill-Decode Disaggregation: Mapping the Cost-Latency Frontier for MoE LLM Serving on H200
TL;DR — Combining prefill/decode disaggregation with tiered KV-cache offloading is analytically antagonistic: disaggregation raises cache hit value but consumes the TTFT slack the slowest tier needs. Empirical validation failed (vLLM init crash), yielding zero performance data.
ThakiCloud AI Research · 2026-07-28 · 📝 Tech blog (KO)
Problem
LLM… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-28-kv-cache-tiering-pd-disagg-cost.Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.details_xformAI__opt-125m-gqa-ub-6-best-for-KV-cache
Dataset Card for Evaluation run of xformAI/opt-125m-gqa-ub-6-best-for-KV-cache
Dataset automatically created during the evaluation run of model xformAI/opt-125m-gqa-ub-6-best-for-KV-cache on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xformAI__opt-125m-gqa-ub-6-best-for-KV-cache.kvcache-quantization-logs-qwen7bdetails_xformAI__facebook-opt-125m-qcqa-ub-6-best-for-KV-cache
Dataset Card for Evaluation run of xformAI/facebook-opt-125m-qcqa-ub-6-best-for-KV-cache
Dataset automatically created during the evaluation run of model xformAI/facebook-opt-125m-qcqa-ub-6-best-for-KV-cache on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xformAI__facebook-opt-125m-qcqa-ub-6-best-for-KV-cache.turboquant-tcq-kv-cache
Closing the Gap: Trellis-Coded Quantization for KV Cache at 2-3 Bits
Authors: buun, Claude (Anthropic)
This dataset accompanies the paper Closing the Gap: Trellis-Coded Quantization for KV Cache at 2-3 Bits. It contains the trained TCQ codebooks, codebook training scripts, and the full paper PDF.
Summary
We present the first application of trellis-coded quantization (TCQ) to KV cache compression in LLM inference. TCQ constrains quantization indices to follow a… See the full description on the dataset page: https://huggingface.co/datasets/spiritbuun/turboquant-tcq-kv-cache.
