leafspark/Meta-Llama-3.1-405B-Instruct-GGUF
2192
Meta-Llama-3.1-405B-Instruct-GGUF

Low bit quantizations of Meta's Llama 3.1 405B Instruct model. Quantized from ollama q4_0 GGUF.
Quantized with llama.cpp b3449
For higher quality quantizations (q4+), please refer to nisten/meta-405b-instruct-cpu-optimized-gguf.
Regarding the smaug-bpe tokenizer, this doesn't make a difference (they are identical). However, if you have concerns you can use the following command to set the llama-bpe tokenizer:
./gguf-py/scripts/gguf_new_metadata.py --pre-tokenizer "llama-bpe" Llama-3.1-405B-Instruct-old.gguf LLama-3.1-405B-Instruct-fixed.ggufimatrix
Generated from Q2_K quant.
imatrix calibration data: groups_merged.txt
