CoolFace
20 results

kvcache

bundaumamtom /kvcache-quantization-logs-qwen7b0 likes77 downloads2mo agoHugging Facest192011 /KVCaches2 likes52 downloads4mo agoHugging Facewangqiyuan2026 /kvcache-bench-results Mingxin KV-Cache Tiered-Storage Benchmark Results Measured results for LLM KV Cache tiered-storage acceleration on 8× AMD Instinct MI308X (ROCm 7.2, vLLM 0.20.1+rocm721, LMCache upstream mainline), model Qwen3-Coder-480B-FP8, published by Mingxin Technology. Device under test: Mingxin FX100 all-flash NVMe-oF array (4-disk RAID0, RoCEv2, 100 GbE) vs local NVMe (PCIe Gen4) vs recompute-only baseline. Files kvcache_bench_results.json — all experiments in structured… See the full description on the dataset page: https://huggingface.co/datasets/wangqiyuan2026/kvcache-bench-results.textn<1K1 likes48 downloads2mo agoHugging Facejprivera44 /kv_cache_logits_exp2textn<1K1 likes43 downloads9mo agoHugging FaceVisionarys /kvcache2 likes42 downloads8mo agoHugging FaceRohithreddybc /kv-cache-compression-mbe Matched-Budget Evaluation (MBE) — KV Cache Compression A standardized reporting protocol for KV cache compression in LLM inference. MBE is not a new task benchmark; it is a thin reporting layer that fixes which models, tasks, and budgets results are reported at, so that numbers from different papers become comparable. Manifest (mbe_manifest.json): the frozen evaluation specification — model suite, task suite (consuming existing benchmarks: LongBench, RULER, SCBench, GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/Rohithreddybc/kv-cache-compression-mbe.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face