batiai/Llama-4-Scout-17B-16E-Instruct-GGUF
0595
Llama 4 Scout 17B-16E-Instruct GGUF — Quantized by BatiAI
<p align="center"> <a href="https://flow.bati.ai"><img src="https://img.shields.io/badge/BatiFlow-on--device%20AI-blue?style=for-the-badge&logo=apple" alt="BatiFlow"></a> <a href="https://ollama.com/batiai/llama4-scout"><img src="https://img.shields.io/badge/Ollama-batiai%2Fllama4--scout-green?style=for-the-badge" alt="Ollama"></a> </p>
imatrix-calibrated GGUF quantizations of meta-llama/Llama-4-Scout-17B-16E-Instruct (109B total / 17B active MoE, 16 experts, multimodal). Quantized directly from official Meta BF16 weights by BatiAI.
Why Llama 4 Scout?
- 109B total / 17B active Mixture-of-Experts (16 experts, top-1 routing) — efficient for size
- Multimodal native: text + vision via
mmproj(image-text-to-text) - Multilingual: 8 official languages + general multilingual
- Native tool calling + extended context
- Meta Llama 4 Community License — commercial-friendly for most cases (see license link)
- Released 2025-04 by Meta
Quick Start
# Q4_K_M (recommended balance, 60GB, M4 Max 128GB ~ M2 Ultra 192GB)
ollama pull batiai/llama4-scout:q4
# IQ3_XXS (smallest, 38GB, M4 Max 64GB+)
ollama pull batiai/llama4-scout:iq3
# Q5_K_M (higher quality, 72GB, M2 Ultra 192GB+)
ollama pull batiai/llama4-scout:q5Available Quantizations
Mac note on `Q3_K_M`: in every model we've benchmarked on Apple Silicon, Q3KM generated slower than Q4KM despite the smaller file — Granite 4.1 (+27%), Gemma 4 26B (+12%), Qwen3.8‑27B (+18%), Qwen3.6‑27B (+8%), on both M4 Max and M4 mini. Metal's Q3K path is limited by dequantization compute rather than bandwidth. We have not measured this particular model's Q3/Q4 pair yet, so treat it as a strong prior, not a measurement: **if Q4K_M fits, take it.** On CUDA the two are effectively tied, so this applies to Macs only.
Multimodal users: also downloadmmproj-*-BF16.gguf(ormmproj-*-Q6_K.gguf) and use withllama-server --mmprojorllama-mtmd-cli.
Hardware Reality Check
How to run
Ollama (text-only)
ollama pull batiai/llama4-scout:q4
ollama run batiai/llama4-scout:q4llama.cpp (text + vision via mmproj)
# Download GGUF + mmproj
hf download batiai/Llama-4-Scout-17B-16E-Instruct-GGUF \
--include "*Q4_K_M*" --include "mmproj-*-Q6_K.gguf" \
--local-dir ./llama4-scout
# Run with vision
llama-mtmd-cli \
-m ./llama4-scout/meta-llama-Llama-4-Scout-17B-16E-Instruct-Q4_K_M.gguf \
--mmproj ./llama4-scout/mmproj-meta-llama-Llama-4-Scout-17B-16E-Instruct-Q6_K.gguf \
--image input.jpg -p "Describe this image."
# Or as a server
llama-server -m ./llama4-scout/meta-llama-Llama-4-Scout-17B-16E-Instruct-Q4_K_M.gguf \
--mmproj ./llama4-scout/mmproj-meta-llama-Llama-4-Scout-17B-16E-Instruct-Q6_K.gguf \
-ngl 99 -c 32768 --port 8080Model details
- Source: meta-llama/Llama-4-Scout-17B-16E-Instruct
- Architecture:
Llama4ForConditionalGeneration— 109B total / 17B active MoE - Experts: 16 routed (top-1 per token) — efficient sparse MoE
- Multimodal: text backbone + vision encoder (mmproj 분리)
- Original precision: BF16
- License: Meta Llama 4 Community License (commercial use OK for most, see link)
BatiAI signing
All GGUFs in this repo carry:
general.author = BatiAIgeneral.url = https://flow.bati.ai
Why BatiAI?
- Quantized directly from official Meta BF16 weights — no re-quantization
- IQ + K-quant variants share the same wikitext-2-raw imatrix recipe as every BatiAI model
- Multimodal mmproj packaged together for one-stop multimodal usage
- Verified on Apple Silicon (M4 Max / M2 Ultra)
License
Inherits Meta Llama 4 Community License. Commercial-friendly for organizations with < 700M MAU. See:
About BatiFlow
BatiFlow — free on-device AI automation for Mac.
<!-- BENCH-START --> Benchmarks coming once Mac M4 Max / M2 Ultra measurements complete. <!-- BENCH-END -->
