CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thaki-AI /daily-paper-2026-07-28-kv-cache-tiering-pd-disagg-cost KV Cache Tiering Meets Prefill-Decode Disaggregation: Mapping the Cost-Latency Frontier for MoE LLM Serving on H200 TL;DR — Combining prefill/decode disaggregation with tiered KV-cache offloading is analytically antagonistic: disaggregation raises cache hit value but consumes the TTFT slack the slowest tier needs. Empirical validation failed (vLLM init crash), yielding zero performance data. ThakiCloud AI Research · 2026-07-28 · 📝 Tech blog (KO) Problem LLM… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-28-kv-cache-tiering-pd-disagg-cost.1 likes163 downloads2mo agoHugging Face02Boxoffice1280 /Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques BoxOffice Verified Seeds This dataset contains the released BoxOffice seed datasets used in the benchmark pipeline described in the accompanying paper. The release includes ten verified seeds: 7 11 13 17 19 23 29 31 47 73 For each seed, we provide: a full JSONL file containing warmup rows plus evaluation rows an eval JSONL file containing only the evaluation rows a manifest JSON file a validation JSON file with directional warmup counts Layout viewer/ normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.tabulartext-generation1K<n<10K3 likes157 downloads5mo agoHugging Face03open-llm-leaderboard-old /details_xformAI__opt-125m-gqa-ub-6-best-for-KV-cache Dataset Card for Evaluation run of xformAI/opt-125m-gqa-ub-6-best-for-KV-cache Dataset automatically created during the evaluation run of model xformAI/opt-125m-gqa-ub-6-best-for-KV-cache on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xformAI__opt-125m-gqa-ub-6-best-for-KV-cache.1 likes106 downloads3y agoHugging Face04bundaumamtom /kvcache-quantization-logs-qwen7b0 likes78 downloads2mo agoHugging Face05open-llm-leaderboard-old /details_xformAI__facebook-opt-125m-qcqa-ub-6-best-for-KV-cache Dataset Card for Evaluation run of xformAI/facebook-opt-125m-qcqa-ub-6-best-for-KV-cache Dataset automatically created during the evaluation run of model xformAI/facebook-opt-125m-qcqa-ub-6-best-for-KV-cache on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xformAI__facebook-opt-125m-qcqa-ub-6-best-for-KV-cache.2 likes71 downloads3y agoHugging Face06spiritbuun /turboquant-tcq-kv-cache Closing the Gap: Trellis-Coded Quantization for KV Cache at 2-3 Bits Authors: buun, Claude (Anthropic) This dataset accompanies the paper Closing the Gap: Trellis-Coded Quantization for KV Cache at 2-3 Bits. It contains the trained TCQ codebooks, codebook training scripts, and the full paper PDF. Summary We present the first application of trellis-coded quantization (TCQ) to KV cache compression in LLM inference. TCQ constrains quantization indices to follow a… See the full description on the dataset page: https://huggingface.co/datasets/spiritbuun/turboquant-tcq-kv-cache.11 likes70 downloads6mo agoHugging Face07open-llm-leaderboard-old /details_saarvajanik__facebook-opt-6.7b-gqa-ub-16-best-for-KV-cache Dataset Card for Evaluation run of saarvajanik/facebook-opt-6.7b-gqa-ub-16-best-for-KV-cache Dataset automatically created during the evaluation run of model saarvajanik/facebook-opt-6.7b-gqa-ub-16-best-for-KV-cache on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_saarvajanik__facebook-opt-6.7b-gqa-ub-16-best-for-KV-cache.2 likes69 downloads3y agoHugging Face08open-llm-leaderboard-old /details_saarvajanik__facebook-opt-6.7b-qcqa-ub-16-best-for-KV-cache Dataset Card for Evaluation run of saarvajanik/facebook-opt-6.7b-qcqa-ub-16-best-for-KV-cache Dataset automatically created during the evaluation run of model saarvajanik/facebook-opt-6.7b-qcqa-ub-16-best-for-KV-cache on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_saarvajanik__facebook-opt-6.7b-qcqa-ub-16-best-for-KV-cache.2 likes68 downloads3y agoHugging Face09vimalnakrani /stillwarm-kv-cache-artifact A downloadable KV-cache save file — with the honest math One llama-server slot save: the first 8,192 Llama-tokens of Frankenstein (public domain), prefilled by Qwen2.5-7B-Instruct Q4_K_M (Apache-2.0 model — chosen over Llama specifically for artifact licensing) and saved with a stillwarm sidecar. This file is USELESS unless your setup matches the sidecar exactly: field value llama.cpp build b9871 (ef2d770117db45b05aa7ecd1b0acca36370c5470) — advisory: ±5 weeks measured… See the full description on the dataset page: https://huggingface.co/datasets/vimalnakrani/stillwarm-kv-cache-artifact.tabularn<1K1 likes65 downloads3mo agoHugging Face10Nathan-Maine /dgx-spark-kv-cache-benchmark KV Cache Quantization on NVIDIA DGX Spark GB10 Corrected benchmarks (v3, April 2026) — KV cache quantization behavior on the NVIDIA DGX Spark's GB10 Grace Blackwell unified memory architecture. Author: Nathan Maine Date: March 2026, corrected April 2026 Hardware: NVIDIA DGX Spark (GB10, compute 12.1, 128GB unified memory) Correction Notice: The original v1 benchmarks (March 31) contained methodology errors. Memory was measured via RSS (wrong on unified memory) and some… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/dgx-spark-kv-cache-benchmark.othern<1K2 likes61 downloads4mo agoHugging Face11st192011 /KVCaches2 likes51 downloads4mo agoHugging Face12wangqiyuan2026 /kvcache-bench-results Mingxin KV-Cache Tiered-Storage Benchmark Results Measured results for LLM KV Cache tiered-storage acceleration on 8× AMD Instinct MI308X (ROCm 7.2, vLLM 0.20.1+rocm721, LMCache upstream mainline), model Qwen3-Coder-480B-FP8, published by Mingxin Technology. Device under test: Mingxin FX100 all-flash NVMe-oF array (4-disk RAID0, RoCEv2, 100 GbE) vs local NVMe (PCIe Gen4) vs recompute-only baseline. Files kvcache_bench_results.json — all experiments in structured… See the full description on the dataset page: https://huggingface.co/datasets/wangqiyuan2026/kvcache-bench-results.textn<1K1 likes48 downloads2mo agoHugging Face13jprivera44 /kv_cache_logits_exp2textn<1K1 likes43 downloads9mo agoHugging Face14Anticloud /camus-10-kv-cache-quantization We attempted KV cache quantization to Q4 — and documented why it fails on the qwen2vl architecture. KV Cache Quantization Attempt: type_k/type_v on qwen2vl Architecture The Problem The KV cache consumes significant memory bandwidth during autoregressive generation. On CPU-bound systems, memory bandwidth is the primary bottleneck. Quantizing the KV cache from FP16 to Q4 theoretically halves memory bandwidth requirements. What We Built We attempted… See the full description on the dataset page: https://huggingface.co/datasets/Anticloud/camus-10-kv-cache-quantization.1 likes43 downloads3mo agoHugging Face15Visionarys /kvcache2 likes42 downloads8mo agoHugging Face16memoriant /dgx-spark-kv-cache-benchmark KV Cache Quantization on NVIDIA DGX Spark GB10 Corrected benchmarks (v3, April 2026) — KV cache quantization behavior on the NVIDIA DGX Spark's GB10 Grace Blackwell unified memory architecture. Author: Nathan Maine, Memoriant Inc. Date: March 2026, corrected April 2026 Hardware: NVIDIA DGX Spark (GB10, compute 12.1, 128GB unified memory) Correction Notice: The original v1 benchmarks (March 31) contained methodology errors. Memory was measured via RSS (wrong on unified memory) and… See the full description on the dataset page: https://huggingface.co/datasets/memoriant/dgx-spark-kv-cache-benchmark.othern<1K0 likes42 downloads6mo agoHugging Face17kleinnner /camus-10-kv-cache-quantization We attempted KV cache quantization to Q4 — and documented why it fails on the qwen2vl architecture. KV Cache Quantization Attempt: type_k/type_v on qwen2vl Architecture The Problem The KV cache consumes significant memory bandwidth during autoregressive generation. On CPU-bound systems, memory bandwidth is the primary bottleneck. Quantizing the KV cache from FP16 to Q4 theoretically halves memory bandwidth requirements. What We Built We attempted… See the full description on the dataset page: https://huggingface.co/datasets/kleinnner/camus-10-kv-cache-quantization.2 likes35 downloads3mo agoHugging Face18Rohithreddybc /kv-cache-compression-mbe Matched-Budget Evaluation (MBE) — KV Cache Compression A standardized reporting protocol for KV cache compression in LLM inference. MBE is not a new task benchmark; it is a thin reporting layer that fixes which models, tasks, and budgets results are reported at, so that numbers from different papers become comparable. Manifest (mbe_manifest.json): the frozen evaluation specification — model suite, task suite (consuming existing benchmarks: LongBench, RULER, SCBench, GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/Rohithreddybc/kv-cache-compression-mbe.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face19h4shk4t /fast-kv-compaction-cache0 likes3 downloads5mo agoHugging Face20ergt2025 /kvcache_offloadgated2 likes1 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.