models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
TinyLlama-1.1B-compressed-tensors-kv-cache-schemeQwen3.8-27B-NVFP4-IRIS-Cache-ObjectScriptopt-125m-gqa-ub-6-best-for-KV-cachefacebook-opt-6.7b-gqa-ub-16-best-for-KV-cachefacebook-opt-125m-qcqa-ub-6-best-for-KV-cachefacebook-opt-6.7b-qcqa-ub-16-best-for-KV-cacheQwen3.5-122B-A10B-FP8-CacheReadyjetson-hybrid-flat-cache-0.5bGLM-4.5-Air-NVFP4-KV-cache-FP8GLM-4.5-Air-NVFP4-KV-cache-BF16GLM-4.7-NVFP4-KV-cache-BF16phi-tiny-moe-cache-rewardGLM-4.7-NVFP4-KV-cache-FP8Qwen3.5-122B-A10B-CacheReadyGLM-4.5-Air-NVFP4-KV-cache-NVFP4vicuna-13b-v1.5-no-cachevicuna-13b-v1.5-16k-no-cacheolmoe-cache-reward0604_key_cache_qwen3_8b_newFix-Strict_Darpo-cache-adapter-3k0604_key_cache_qwen3_8bjetson-flat-cache-cpt-adaptergemma-2-2b-lean-expert-optimized-cache-enabledzay-qwen7b-codem-4gpu-v10-cachedgpt2_coreml_kv_cache_try1test-cacheyTinySwallow-1.5B-Instruct-q4f32_1-MLC-tensor-cache-mirror
