models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
TinyLlama-1.1B-Chat-v1.0-kvcache-fp8-tensorTinyLlama-1.1B-Chat-v1.0-kvcache-fp8-attn_headkv_cache_fp8-e2ekv_cache_gptq_tinyllama-e2eTinyLlama-1.1B-compressed-tensors-kv-cache-schemeopt-125m-gqa-ub-6-best-for-KV-cachefacebook-opt-6.7b-gqa-ub-16-best-for-KV-cachefacebook-opt-125m-qcqa-ub-6-best-for-KV-cachefacebook-opt-6.7b-qcqa-ub-16-best-for-KV-cachepixtral-12b-FP8-dynamic-FP8-KV-cacheQwen3-30BA3B-GGUFGLM-4.5-Air-NVFP4-KV-cache-FP8GLM-4.5-Air-NVFP4-KV-cache-BF16Kimi-K2-Instruct-GGUFKimi-K2-Instruct-0905-GGUFGLM-4.7-NVFP4-KV-cache-BF16GLM-4.7-NVFP4-KV-cache-FP8manga-ocr-kvcache-tfliteGLM-4.5-Air-NVFP4-KV-cache-NVFP4llama32-1b-kvcache-coreml-macosLlama-3.1-8B-Instruct-KV-Cache-FP8sampling_with_kvcacheKimi-K2-Thinking-CPU-weightPhi-3-mini-4k-instruct-kv_cache_default_phi3-e2eTinyLlama-1.1B-Chat-v1.0-kv_cache_default_tinyllama-e2ellama32-1b-kvcache-coreml-iosqmd-query-expansion-1.7B-ONNX-kvcache-fp32sampling_with_kvcache_hf_helpersLlama-3-70B-Instruct-awq-int8-kv-cache-trt-llm-compiledMeta-Llama-3-8B-Instruct-FP8-channel-output-activation-kv_cache-qkv_proj
