CoolFace
20 results

kv_cache

thaki-AI /daily-paper-2026-07-28-kv-cache-tiering-pd-disagg-cost KV Cache Tiering Meets Prefill-Decode Disaggregation: Mapping the Cost-Latency Frontier for MoE LLM Serving on H200 TL;DR — Combining prefill/decode disaggregation with tiered KV-cache offloading is analytically antagonistic: disaggregation raises cache hit value but consumes the TTFT slack the slowest tier needs. Empirical validation failed (vLLM init crash), yielding zero performance data. ThakiCloud AI Research · 2026-07-28 · 📝 Tech blog (KO) Problem LLM… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-28-kv-cache-tiering-pd-disagg-cost.1 likes163 downloads2mo agoHugging FaceBoxoffice1280 /Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques BoxOffice Verified Seeds This dataset contains the released BoxOffice seed datasets used in the benchmark pipeline described in the accompanying paper. The release includes ten verified seeds: 7 11 13 17 19 23 29 31 47 73 For each seed, we provide: a full JSONL file containing warmup rows plus evaluation rows an eval JSONL file containing only the evaluation rows a manifest JSON file a validation JSON file with directional warmup counts Layout viewer/ normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.tabulartext-generation1K<n<10K3 likes157 downloads5mo agoHugging Faceopen-llm-leaderboard-old /details_xformAI__opt-125m-gqa-ub-6-best-for-KV-cache Dataset Card for Evaluation run of xformAI/opt-125m-gqa-ub-6-best-for-KV-cache Dataset automatically created during the evaluation run of model xformAI/opt-125m-gqa-ub-6-best-for-KV-cache on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xformAI__opt-125m-gqa-ub-6-best-for-KV-cache.1 likes106 downloads3y agoHugging Facebundaumamtom /kvcache-quantization-logs-qwen7b0 likes78 downloads2mo agoHugging Faceopen-llm-leaderboard-old /details_xformAI__facebook-opt-125m-qcqa-ub-6-best-for-KV-cache Dataset Card for Evaluation run of xformAI/facebook-opt-125m-qcqa-ub-6-best-for-KV-cache Dataset automatically created during the evaluation run of model xformAI/facebook-opt-125m-qcqa-ub-6-best-for-KV-cache on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xformAI__facebook-opt-125m-qcqa-ub-6-best-for-KV-cache.2 likes71 downloads3y agoHugging Facespiritbuun /turboquant-tcq-kv-cache Closing the Gap: Trellis-Coded Quantization for KV Cache at 2-3 Bits Authors: buun, Claude (Anthropic) This dataset accompanies the paper Closing the Gap: Trellis-Coded Quantization for KV Cache at 2-3 Bits. It contains the trained TCQ codebooks, codebook training scripts, and the full paper PDF. Summary We present the first application of trellis-coded quantization (TCQ) to KV cache compression in LLM inference. TCQ constrains quantization indices to follow a… See the full description on the dataset page: https://huggingface.co/datasets/spiritbuun/turboquant-tcq-kv-cache.11 likes70 downloads6mo agoHugging Face