majentik/Qwen3-Embedding-0.6B-GGUF-IQ4_XS
162
Qwen3-Embedding-0.6B GGUF IQ4_XS
llama.cpp GGUF IQ4_XS quantization of Qwen/Qwen3-Embedding-0.6B.
- Produced with:
llama-quantize(upstream llama.cpp, April 2026 build) - BF16 source converted via
convert_hf_to_gguf.pyfrom the fresh llama.cpp tree - Quant type: IQ4_XS
- File size: 352 MB
Quickstart
llama-embedding -m qwen3-emb-0.6b-IQ4_XS.gguf \
-p "What is the capital of France?"Or via llama-cpp-python:
from llama_cpp import Llama
llm = Llama(model_path="qwen3-emb-0.6b-IQ4_XS.gguf", embedding=True)
vec = llm.embed("What is the capital of France?")License
Apache 2.0 — inherited from the upstream base model.
See also
- Base: Qwen/Qwen3-Embedding-0.6B
- Garden hub: majentik/garden
- llama.cpp: https://github.com/ggml-org/llama.cpp
