CoolFace
Datasetpublic

thaki-AI/daily-paper-2026-07-28-kv-cache-tiering-pd-disagg-cost

KV Cache Tiering Meets Prefill-Decode Disaggregation: Mapping the Cost-Latency Frontier for MoE LLM Serving on H200 TL;DR — Combining prefill/decode disaggregation with tiered KV-cache offloading is analytically antagonistic: disaggregation raises cache hit value but consumes the TTFT slack the slowest tier needs. Empirical validation failed (vLLM init crash), yielding zero performance data. ThakiCloud AI Research · 2026-07-28 · 📝 Tech blog (KO) Problem LLM… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-28-kv-cache-tiering-pd-disagg-cost.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes144downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face