majentik/Qwen3-Embedding-0.6B-GGUF-Q4_K_M
045
Qwen3-Embedding-0.6B GGUF Q4KM
llama.cpp GGUF Q4KM quantization of Qwen/Qwen3-Embedding-0.6B.
- Produced with:
llama-quantize(upstream llama.cpp, April 2026 build) - BF16 source converted via
convert_hf_to_gguf.pyfrom the fresh llama.cpp tree - Quant type: Q4_K_M
- File size: 378 MB
Quickstart
llama-embedding -m qwen3-emb-0.6b-Q4_K_M.gguf \
-p "What is the capital of France?"Or via llama-cpp-python:
from llama_cpp import Llama
llm = Llama(model_path="qwen3-emb-0.6b-Q4_K_M.gguf", embedding=True)
vec = llm.embed("What is the capital of France?")License
Apache 2.0 — inherited from the upstream base model.
See also
- Base: Qwen/Qwen3-Embedding-0.6B
- Garden hub: majentik/garden
- llama.cpp: https://github.com/ggml-org/llama.cpp
