abhimittal/kv-cache-budget-simulator
0
KV Cache Budget Simulator
Exact KV-cache memory math for LLM serving — GQA/MQA/MLA, KV quantization, PagedAttention fragmentation, OOM walls, and the latency–throughput Pareto frontier at a fixed memory budget. Pure client-side, no GPU needed.
Source: https://github.com/aabhimittal/Attention-trace-differ-for-speculative-decoding
