batiai/Qwen3-Reranker-0.6B-GGUF
Qwen3-Reranker-0.6B GGUF — Quantized by BatiAI
<p align="center"> <a href="https://flow.bati.ai"><img src="https://img.shields.io/badge/BatiFlow-RAG%20on%20Mac-blue?style=for-the-badge&logo=apple" alt="BatiFlow"></a> <a href="https://huggingface.co/Qwen/Qwen3-Reranker-0.6B"><img src="https://img.shields.io/badge/Upstream-Qwen3--Reranker--0.6B-orange?style=for-the-badge" alt="Upstream"></a> </p>
GGUF quantizations of Qwen/Qwen3-Reranker-0.6B — the most-downloaded open-source reranker of 2026 (1.39 M downloads on HF). Part of BatiAI's on-device RAG stack for BatiFlow.
What is a reranker?
RAG pipeline: embedding (coarse retrieve) → reranker (precise scoring) → LLM (answer).
A reranker takes (query, candidate_document) and returns a relevance score. It's the "second pass" after vector search — turns "probably relevant" candidates into an ordered top-K that the LLM can use confidently.
Quick Start (llama.cpp)
./llama-cli -m Qwen3-Reranker-0.6B-Q6_K.gguf \
--chat-template-file chat-template.jinja \
-p "<query>weather in Seoul</query><doc>Seoul had rain yesterday</doc>"For production, integrate via the llama.cpp API (see Qwen3-Reranker usage).
Note: Ollama doesn't have a native reranker endpoint yet, so this GGUF is intended for direct llama.cpp integration or tools like LangChain / LlamaIndex.
Available Quantizations
Small models don't benefit much from aggressive quantization (IQ3/IQ4 degrades ranking quality). Q6_K is the sweet spot.
Quality Verification (measured)
Ran 40 (query, positive, negative) triples — 20 EN + 20 KO — twice:
- Easy — off-topic negatives (e.g. "Eiffel Tower" as negative for "gradient descent")
- Hard — topically-close negatives (e.g. "backpropagation" as negative for "gradient descent")
Pearson correlation of scores Q6_K ↔ Q8_0: r = 0.998 on hard test → quantization drift is under measurement noise. Q6_K is safe.
Full bench reports in reports/rerank-quality-* of the pipeline repo. Reproducible with `scripts/bench-rerank-quality.sh`.
Why Qwen3-Reranker?
- SOTA among open rerankers — top of MTEB reranking benchmarks
- Multilingual — English / Korean / Japanese / Chinese
- Tiny footprint — 0.6B parameters, fits in 1 GB RAM
- Apache 2.0 — commercial-friendly
Why BatiAI?
- Quantized directly from Alibaba's BF16 safetensors — no intermediate GGUF
- BatiAI-signed —
general.author: BatiAI,general.url: https://flow.bati.ai - Part of a full on-device RAG stack (chat LLM + reranker + embedding) — see the batiai HF profile
Technical Details
- Original Model: Qwen/Qwen3-Reranker-0.6B
- Architecture: Qwen3 Causal LM (used as cross-encoder scorer)
- Parameters: 596 M
- Context: 32 K
- License: Apache 2.0
- Quantized with: llama.cpp build
bafae2765
About BatiAI's RAG Stack
License
Mirrors upstream Qwen Apache 2.0. Commercial use permitted.
