CoolFace
Modelpublic

majentik/Qwen3-Embedding-0.6B-MLX-4bit

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes40downloads
Model Card

Qwen3-Embedding-0.6B MLX 4-bit

MLX 4-bit quantization of Qwen/Qwen3-Embedding-0.6B, produced with mlx-embeddings on Apple Silicon.

What is this?

Qwen3-Embedding is a decoder-only LLM-style text embedding model from the Qwen3 family, using last-token pooling to produce dense vector representations. It scores near the top of MMTEB multilingual benchmarks while retaining Apache-2.0 licensing.

Quantization

  • —Method: MLX affine quantization (mlx_embeddings.convert), group_size=64
  • —Bits per weight: 4
  • —Output size: 333 MB (vs ~1.2 GB for bf16 source)

Quickstart

python
from mlx_embeddings import load

model, tokenizer = load("majentik/Qwen3-Embedding-0.6B-MLX-4bit")

inputs = tokenizer(
    ["What is the capital of France?", "Paris is the capital of France."],
    padding=True, truncation=True, return_tensors="mlx"
)
outputs = model(inputs["input_ids"], attention_mask=inputs["attention_mask"])
embeddings = outputs.text_embeds  # already L2-normalised, shape [batch, dim]

For sentence similarity:

python
import mlx.core as mx

e = embeddings
scores = (e[0] @ e[1:].T).tolist()
print(scores)

Model Specifications

PropertyValue
Base ModelQwen/Qwen3-Embedding-0.6B
ArchitectureDecoder-only (Qwen3ForCausalLM) with last-token pooling
Parameters0.6B (596M) (pre-quantization)
Context Length32K
Embedding Dim1024
BF16 Size~1.2 GB
Licenseapache-2.0
Languages100+ (multilingual)

License

Apache 2.0 — inherited from the upstream Qwen3-Embedding model. Free for research and commercial use.

See also