CoolFace
Modelpublic

dispatchAI/Llama-3.2-1B-Instruct-Q6-mobile

sourceHugging Facellama3.2updated 3mo agoView on Hugging Face
0likes29downloads
Model Card

Llama 3.2 1B Instruct - Q6 Mobile (GGUF) - RECOMMENDED

The balanced choice between tiny Q4 (~767MB) and larger Q8 (~1.3GB). Best quality-to-size ratio for most deployments.

PropertyValue
Parameters1.23 billion
QuantizationQ6_K (6-bit)
Size~1.02 GB
Speed~25 tok/s (S20 FE CPU)
Quality Retention~97%

This is the recommended default variant when unsure which to choose.