models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
MiniMax-M3-EAGLE3-GQAtiny-random-LlamaForCausalLM-GQATinyStories-LLaMA2-25M-256h-4l-GQAsmol_llama-220M-GQAsmol_llama-101M-GQAEYE-Llama_gqasmol_llama-220M-GQA-32k-theta-sft-limarpTinyStories-LLaMA2-42.5M-512h-4l-GQATinyStories-LLaMA2-20M-256h-4l-GQAllama-test-gqa-with-better-transformerMixtral-GQA-400m-v2tiny-random-LlamaForCausalLM-GQAMiniMax-M3-EAGLE3-GQA-NVFP4TinyStories-LLaMA2-20M-256h-4l-GQATinyStories-LLaMA2-20M-256h-4l-GQATinyStories-LLaMA2-20M-256h-4l-GQATinyStories-LLaMA2-20M-256h-4l-GQATinyStories-LLaMA2-20M-256h-4l-GQATinyStories-LLaMA2-20M-256h-4l-GQA-xiangziTinyStories-LLaMA2-20M-256h-4l-GQATinyStories-LLaMA2-20M-256h-4l-GQAllama2_xs_233m_GQA-llama-1028-interleaved-deduped-v1-tb-interleaved-deduped-1028-0919opt-125m-gqa-ub-6-best-for-KV-cachefacebook-opt-6.7b-gqa-ub-16-best-for-KV-cachesmol_llama-101M-GQA-pythonsmol_llama-220M-GQA-fineweb_edusmol_llama-220M-GQA-32k-theta-sftsmol_llama-220M-GQA-32k-linearsmol_llama-220M-GQA-32k-thetaNanoLlama-GQA-L10-A32_KV8-v13-KI
