CoolFace
Modelpublic

groxaxo/Nemotron-3-Embed-1B-oQ5-MLX

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
1likes68downloads
Model Card

Nemotron-3-Embed-1B-oQ5-MLX

<!-- polished-overview:start -->

Overview

Nemotron-3-Embed-1B-oQ5-MLX is an MLX-formatted checkpoint optimized for Apple silicon, published by `groxaxo`. It is intended for open-source evaluation, reproducible experimentation, and compatible local or hosted inference workflows. The wording below is deliberately limited to what can be verified from this repository's metadata and artifacts.

At a glance

FieldDetails
FormatMLX
Source / base`nvidia/Nemotron-3-Embed-1B-BF16`
Intended taskfeature-extraction
Licenseother

What is included

  • —*.safetensors (1 file)
  • —config.json
  • —tokenizer.json
  • —tokenizer_config.json
  • —Additional configuration, tokenizer, processor, or shard files (10 visible artifacts total)

Quick start

MLX

Use an up-to-date MLX-compatible runtime on Apple silicon and point it at this repository:

bash
mlx_lm.generate --model groxaxo/Nemotron-3-Embed-1B-oQ5-MLX --prompt "Write a concise technical summary."

Embedding and audio repositories may require the task-specific MLX package documented by the upstream project.

Compatibility and responsible use

  • —Use a runtime that explicitly supports this format, architecture, and modality.
  • —Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • —Review the source model card and license before redistribution or deployment.
  • —Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • —Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws. <!-- polished-overview:end -->

An Apple-Silicon-native OMLX oQ5 conversion of `nvidia/Nemotron-3-Embed-1B-BF16`.

791 MiB complete checkout · 775 MiB model shard · 5.70 effective bits/weight · 2,048-dimensional normalized embeddings · MLX

This checkpoint was quantized locally on a 24 GiB Apple-silicon Mac using `jundot/omlx`. OMLX measured all 16 encoder layers with 128 calibration samples × 256 tokens, selected a base 5-bit affine plan with group size 64, and promoted 15 sensitive projections to 6 or 8 bits.

The model preserves the source embedding contract:

  • —bidirectional attention (is_causal: false)
  • —average/mean pooling
  • —L2-normalized 2,048-dimensional output
  • —query: prefix for queries
  • —passage: prefix for documents

oQ5 versus the existing AWQ build

The companion CUDA model is `groxaxo/Nemotron-3-Embed-1B-AWQ-W4A16`. They target different machines and serving stacks.

PropertyThis OMLX oQ5 checkpointExisting AWQ W4A16 checkpoint
Primary hardwareApple SiliconNVIDIA CUDA, especially Ampere/Ada
RuntimeMLX / mlx-lm encoder pathvLLM with Marlin
Weight formatMLX affine mixed precisioncompressed-tensors AWQ
Base weight width5 bit4 bit
Effective plan5.70 bpwW4A16 Linear layers
Group size6432
Higher-precision protection15 sensitivity-selected projections at 6/8 bitBF16 embedding table and norms
Model shard812,820,814 bytesapproximately 994 MB
Complete local checkoutapproximately 791 MiBapproximately 1.0 GB
Best useLocal embedding/RAG on a MacHigh-throughput embedding service on NVIDIA GPUs

The formats are not interchangeable. Use this checkpoint with MLX on Apple Silicon. Use the AWQ checkpoint with vLLM/Marlin on NVIDIA hardware.

Direct head-to-head retrieval test

Both checkpoints were evaluated against the same corpus from `groxaxo/mcp-chrome-patched`:

  • —10 repository-specific queries
  • —12 passages read directly from tracked documentation and package.json
  • —one labelled relevant passage per query
  • —identical query: and passage: role formatting
  • —AWQ evaluated through its live vLLM endpoint
  • —oQ5 evaluated directly through MLX
MetricAWQ W4A16 / vLLMOMLX oQ5 / MLX
Top-1 accuracy90%90%
Mean reciprocal rank0.9333330.933333
Mean relevant-score margin0.1912030.191718
Output dimensions2,0482,048

Cross-format agreement:

  • —retrieval-score Pearson correlation: 0.999085
  • —mean absolute score delta: 0.005904
  • —mean same-input AWQ↔oQ5 vector cosine: 0.994057
  • —minimum same-input vector cosine: 0.992844
  • —identical top-three ordering for all ten queries

Both models made the same one miss: for a query asking which package script builds the shared types, both ranked two prose passages about shared architecture/build flow above the shorter package.json passage containing the exact script. This points to query/corpus ambiguity, not a quantization-specific regression.

Important benchmark boundary

This small, domain-specific test shows that the two quantizations behave almost identically on the Chrome MCP repository. It does not prove universal quality equivalence.

The AWQ repository includes the broader held-out English PAQ and Spanish MIRACL-ES evaluations. This oQ5 release has currently been validated with the Chrome MCP comparison above plus direct BF16 fidelity checks; it has not yet been rerun on the AWQ repository's 512-pair English and Spanish suites.

BF16 fidelity check

Three role-prefixed inputs—a query, relevant passage, and unrelated passage—were embedded with both the official BF16 source and this oQ5 checkpoint.

InputBF16↔oQ5 vector cosine
Query0.995496
Relevant passage0.995863
Unrelated passage0.994637

Semantic scores:

ComparisonBF16oQ5
Query ↔ relevant passage0.6805750.687804
Query ↔ unrelated passage-0.024595-0.012204

All outputs were 2,048-dimensional and L2-normalized.

Quick start on Apple Silicon

Create an environment with MLX-LM and download the checkpoint:

bash
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade mlx mlx-lm transformers huggingface_hub

hf download groxaxo/Nemotron-3-Embed-1B-oQ5-MLX \
  --local-dir ./Nemotron-3-Embed-1B-oQ5-MLX

Run the included bidirectional encoder example:

bash
python ./Nemotron-3-Embed-1B-oQ5-MLX/embed_mlx.py \
  --model ./Nemotron-3-Embed-1B-oQ5-MLX \
  --query "Which city is known as the City of Sails?" \
  --passage "Auckland is widely known as the City of Sails."

The script prints the vector dimensions, L2 norms, and query-to-passage cosine similarity.

Why the included encoder path matters

This is a bidirectional embedding checkpoint, not a causal language model. Calling the normal mlx-lm generation model interface would construct a causal mask and would not reproduce the intended embeddings. embed_mlx.py loads the bare Ministral3 language-model backbone, runs each block without a causal mask, mean-pools token states, and L2-normalizes the output.

OMLX support for correctly calibrating this checkpoint was contributed in `jundot/omlx#2410`. Until that PR is merged, the tested implementation is also available on the `groxaxo/omlx` feature branch.

Quantization details

SettingValue
OMLX source revisioncfa2f445de84fe54ad2648ee8b2c9aec27fb1a17
Base model revisiona5e0f804b9e90a1ca6784ecbf6e41595774fc834
OMLX leveloQ5
Base quantization5-bit affine
Group size64
Working dtypebfloat16
Calibration128 samples × 256 tokens
Effective plan5.70 bpw
Sensitivity boostseight 8-bit projections, seven 6-bit projections
Safetensors SHA-2561880be31481999ab2953940544fe6cfcff6637c5a4b4e39661fb8eb02d24ed29

The complete OMLX tests/test_oq.py suite passed: 324 tests.

Included files

FilePurpose
model.safetensorsQuantized MLX weights
config.jsonBidirectional Ministral3 configuration and per-layer quantization plan
tokenizer.json, tokenizer_config.jsonSource tokenizer
embed_mlx.pyStandalone MLX embedding example
QUANTIZATION_REPORT.mdConversion provenance and BF16 validation
CHROME_MCP_COMPARISON.mdDetailed AWQ-versus-oQ5 retrieval comparison
LICENSE, NOTICE, THIRD_PARTY_NOTICES.mdFiles preserved from NVIDIA's source repository

License

This derivative inherits the `OpenMDW-1.1` license from `nvidia/Nemotron-3-Embed-1B-BF16`. Review the included LICENSE, NOTICE, and THIRD_PARTY_NOTICES.md files.

Citation

bibtex
@misc{nvidia_nemotron_3_embed_1b,
  title  = {Nemotron-3-Embed-1B-BF16},
  author = {NVIDIA Corporation},
  url    = {https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16},
  year   = {2026}
}

@misc{groxaxo_nemotron_3_embed_omlx_oq5,
  title  = {Nemotron-3-Embed-1B-oQ5-MLX},
  author = {Facu Vlad J},
  url    = {https://huggingface.co/groxaxo/Nemotron-3-Embed-1B-oQ5-MLX},
  year   = {2026},
  note   = {OMLX oQ5 conversion for Apple Silicon}
}