CoolFace
Modelpublic

BabaK07/Qwen2.5-7B-Instruct-1M-GGUF

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes583downloads
Model Card

Qwen2.5 7B Instruct 1M by Qwen

Model creator: Qwen<br> Original model: Qwen2.5-7B-Instruct-1M<br>

Prompt format

<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant

Download a file (not the whole branch) from below:

FilenameQuant typeFile SizeSplitDescription
Qwen2.5-7B-Instruct-1M-F32.gguff3230.5 GBfalseFull F32 weights.
Qwen2.5-7B-Instruct-1M-F16.gguff1615.24 GBfalseFull F16 weights.
Qwen2.5-7B-Instruct-1M-Q8_0.ggufQ8_08.10 GBfalseExtremely high quality, generally unneeded but max available quant.
Qwen2.5-7B-Instruct-1M-Q6_K.ggufQ6_K6.25 GBfalseVery high quality, near perfect, recommended.
Qwen2.5-7B-Instruct-1M-Q5_K_M.ggufQ5KM5.44 GBfalseHigh quality, recommended.
Qwen2.5-7B-Instruct-1M-Q5_K_S.ggufQ5KS5.32 GBfalseHigh quality, recommended.
Qwen2.5-7B-Instruct-1M-Q4_1.ggufQ4_14.87 GBfalseLegacy format, similar performance to Q4KS but with improved tokens/watt on Apple silicon.
Qwen2.5-7B-Instruct-1M-Q4_K_M.ggufQ4KM4.68 GBfalseGood quality, default size for most use cases, recommended.
Qwen2.5-7B-Instruct-1M-Q4_K_S.ggufQ4KS4.46 GBfalseSlightly lower quality with more space savings, recommended.
Qwen2.5-7B-Instruct-1M-Q4_0.ggufQ4_04.43 GBfalseLegacy format, offers online repacking for ARM and AVX CPU inference.
Qwen2.5-7B-Instruct-1M-Q3_K_L.ggufQ3KL4.09 GBfalseLower quality but usable, good for low RAM availability.
Qwen2.5-7B-Instruct-1M-Q3_K_M.ggufQ3KM3.81 GBfalseLow quality.
Qwen2.5-7B-Instruct-1M-Q3_K_S.ggufQ3KS3.49 GBfalseLow quality, not recommended.
Qwen2.5-7B-Instruct-1M-Q2_K.ggufQ2_K3.02 GBfalseVery low quality but surprisingly usable.

Technical Details

Supports a context length of up to 1M tokens.

Significantly improved performance in handling long-context tasks while maintaining its capability in short tasks.

Accuracy degradation may occur for sequences exceeding 262,144 tokens until improved support is added.

For more information, check their blog here.

Downloading using huggingface-cli

<details> <summary>Click to view download instructions</summary>

First, make sure you have hugginface-cli installed:

pip install -U "huggingface_hub[cli]"

Then, you can target the specific file you want:

huggingface-cli download BabaK07/Qwen2.5-7b-Instruct-1M-Q4_K_M-gguf --include "Qwen2.5-7b-Instruct-1M-Q4_K_M.gguf" --local-dir ./

</details>

Special thanks

๐Ÿ™ Special thanks to Georgi Gerganov and the whole team working on llama.cpp for making all of this possible.