CoolFace
Modelpublic

ncorder/llama-embed-nemotron-8b-mlx-fp16

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes29downloads
Model Card

ncorder/llama-embed-nemotron-8b-mlx-fp16

MLX conversion of `nvidia/llama-embed-nemotron-8b` — the #1 embedding model under 20B parameters on the MTEB leaderboard, outperforming models 3x its size.

  • —Parameters: 7.5B
  • —Quantization: None (float16)
  • —Model size: 15 GB
  • —Architecture: Llama-3.1-8B with bidirectional attention
  • —Embedding dimension: 4096
  • —Max sequence length: 32,768 tokens
  • —Converted with: mlx-embeddings

Full float16 precision. Near-identical to the original bfloat16 weights.

All variants

VariantSizeRelevant ↑Irrelevant ↓Margin ↑
fp1615 GB0.37630.05790.3184
8-bit7.5 GB0.37800.05830.3197
**4-bit**4.0 GB0.38260.07830.3043
2-bit2.4 GB0.47990.28730.1926

Reference (original bf16 PyTorch): relevant=0.3771, irrelevant=0.0581, margin=0.3190

Usage

bash
pip install mlx-embeddings
python
from mlx_embeddings.utils import load
import mlx.core as mx

model, tokenizer = load("ncorder/llama-embed-nemotron-8b-mlx-fp16")

query = "Instruct: Given a question, retrieve passages that answer the question\nQuery: How do neural networks learn patterns from examples?"
document = "Deep learning models adjust their weights through backpropagation."

def embed(text):
    inputs = tokenizer(text, return_tensors="np", padding=True)
    out = model(
        mx.array(inputs["input_ids"]),
        mx.array(inputs["attention_mask"])
    )
    return out.text_embeds

q_emb = embed(query)
d_emb = embed(document)
score = (q_emb @ d_emb.T).item()
print(f"Similarity: {score:.4f}")

Query formatting

This model is instruction-aware. For retrieval, prefix queries with:

Instruct: {task_instruction}
Query: {your_query}

Documents are embedded without any prefix.

License

This model inherits the NVIDIA license from the original — research/non-commercial use only. Also subject to the Llama 3.1 Community License.

Credits