groxaxo/Nemotron-3-Embed-1B-oQ5-MLX
Nemotron-3-Embed-1B-oQ5-MLX
<!-- polished-overview:start -->
Overview
Nemotron-3-Embed-1B-oQ5-MLX is an MLX-formatted checkpoint optimized for Apple silicon, published by `groxaxo`. It is intended for open-source evaluation, reproducible experimentation, and compatible local or hosted inference workflows. The wording below is deliberately limited to what can be verified from this repository's metadata and artifacts.
At a glance
What is included
*.safetensors(1 file)config.jsontokenizer.jsontokenizer_config.json- Additional configuration, tokenizer, processor, or shard files (10 visible artifacts total)
Quick start
MLX
Use an up-to-date MLX-compatible runtime on Apple silicon and point it at this repository:
mlx_lm.generate --model groxaxo/Nemotron-3-Embed-1B-oQ5-MLX --prompt "Write a concise technical summary."Embedding and audio repositories may require the task-specific MLX package documented by the upstream project.
Compatibility and responsible use
- Use a runtime that explicitly supports this format, architecture, and modality.
- Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
- Review the source model card and license before redistribution or deployment.
- Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
- Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.
Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.
Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws. <!-- polished-overview:end -->
An Apple-Silicon-native OMLX oQ5 conversion of `nvidia/Nemotron-3-Embed-1B-BF16`.
791 MiB complete checkout · 775 MiB model shard · 5.70 effective bits/weight · 2,048-dimensional normalized embeddings · MLX
This checkpoint was quantized locally on a 24 GiB Apple-silicon Mac using `jundot/omlx`. OMLX measured all 16 encoder layers with 128 calibration samples × 256 tokens, selected a base 5-bit affine plan with group size 64, and promoted 15 sensitive projections to 6 or 8 bits.
The model preserves the source embedding contract:
- bidirectional attention (
is_causal: false) - average/mean pooling
- L2-normalized 2,048-dimensional output
query:prefix for queriespassage:prefix for documents
oQ5 versus the existing AWQ build
The companion CUDA model is `groxaxo/Nemotron-3-Embed-1B-AWQ-W4A16`. They target different machines and serving stacks.
The formats are not interchangeable. Use this checkpoint with MLX on Apple Silicon. Use the AWQ checkpoint with vLLM/Marlin on NVIDIA hardware.
Direct head-to-head retrieval test
Both checkpoints were evaluated against the same corpus from `groxaxo/mcp-chrome-patched`:
- 10 repository-specific queries
- 12 passages read directly from tracked documentation and
package.json - one labelled relevant passage per query
- identical
query:andpassage:role formatting - AWQ evaluated through its live vLLM endpoint
- oQ5 evaluated directly through MLX
Cross-format agreement:
- retrieval-score Pearson correlation: 0.999085
- mean absolute score delta: 0.005904
- mean same-input AWQ↔oQ5 vector cosine: 0.994057
- minimum same-input vector cosine: 0.992844
- identical top-three ordering for all ten queries
Both models made the same one miss: for a query asking which package script builds the shared types, both ranked two prose passages about shared architecture/build flow above the shorter package.json passage containing the exact script. This points to query/corpus ambiguity, not a quantization-specific regression.
Important benchmark boundary
This small, domain-specific test shows that the two quantizations behave almost identically on the Chrome MCP repository. It does not prove universal quality equivalence.
The AWQ repository includes the broader held-out English PAQ and Spanish MIRACL-ES evaluations. This oQ5 release has currently been validated with the Chrome MCP comparison above plus direct BF16 fidelity checks; it has not yet been rerun on the AWQ repository's 512-pair English and Spanish suites.
BF16 fidelity check
Three role-prefixed inputs—a query, relevant passage, and unrelated passage—were embedded with both the official BF16 source and this oQ5 checkpoint.
Semantic scores:
All outputs were 2,048-dimensional and L2-normalized.
Quick start on Apple Silicon
Create an environment with MLX-LM and download the checkpoint:
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade mlx mlx-lm transformers huggingface_hub
hf download groxaxo/Nemotron-3-Embed-1B-oQ5-MLX \
--local-dir ./Nemotron-3-Embed-1B-oQ5-MLXRun the included bidirectional encoder example:
python ./Nemotron-3-Embed-1B-oQ5-MLX/embed_mlx.py \
--model ./Nemotron-3-Embed-1B-oQ5-MLX \
--query "Which city is known as the City of Sails?" \
--passage "Auckland is widely known as the City of Sails."The script prints the vector dimensions, L2 norms, and query-to-passage cosine similarity.
Why the included encoder path matters
This is a bidirectional embedding checkpoint, not a causal language model. Calling the normal mlx-lm generation model interface would construct a causal mask and would not reproduce the intended embeddings. embed_mlx.py loads the bare Ministral3 language-model backbone, runs each block without a causal mask, mean-pools token states, and L2-normalizes the output.
OMLX support for correctly calibrating this checkpoint was contributed in `jundot/omlx#2410`. Until that PR is merged, the tested implementation is also available on the `groxaxo/omlx` feature branch.
Quantization details
The complete OMLX tests/test_oq.py suite passed: 324 tests.
Included files
License
This derivative inherits the `OpenMDW-1.1` license from `nvidia/Nemotron-3-Embed-1B-BF16`. Review the included LICENSE, NOTICE, and THIRD_PARTY_NOTICES.md files.
Citation
@misc{nvidia_nemotron_3_embed_1b,
title = {Nemotron-3-Embed-1B-BF16},
author = {NVIDIA Corporation},
url = {https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16},
year = {2026}
}
@misc{groxaxo_nemotron_3_embed_omlx_oq5,
title = {Nemotron-3-Embed-1B-oQ5-MLX},
author = {Facu Vlad J},
url = {https://huggingface.co/groxaxo/Nemotron-3-Embed-1B-oQ5-MLX},
year = {2026},
note = {OMLX oQ5 conversion for Apple Silicon}
}