CoolFace
Modelpublic

Spangler3000/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-hybrid8-mtp

sourceHugging Faceotherupdated 15d agoView on Hugging Face
0likes333downloads
Model Card

NVIDIA Nemotron 3.5 Lightning 30B-A3B — MLX hybrid8 (Mamba/attention Q8, experts Q4) with embedded MTP

Status: private validation candidate. Not the published Darkbloom artifact. Not qualified for production. This checkpoint exists to fix a reasoning-off tool-calling regression observed on the published 4-bit listing. Serve it only with a Darkbloom provider that includes the Nemotron stray-</think> parser fix (private d-inference PR #2, commit A); without that fix this artifact emits a bare </think> before its tool call in reasoning-off mode.

What it is

A mixed-precision MLX conversion of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 at source revision a9904d24bcc1d289a1950fa9d2b978c47cf903b9, with the embedded multi-token-prediction (MTP) head, using the structural hybrid8 rule:

tensor groupprecision
Mamba mixer projections (in_proj, out_proj)affine Q8, group 64
attention projections (q/k/v/o)affine Q8, group 64
MoE experts (switch_mlp.fc1/fc2), shared experts, embeddings, lm_headaffine Q4, group 64 (unchanged from the published 4-bit artifact)
router gates, norms, conv1d, A_log, dt_bias, Dunquantized, byte-identical to source
MTP head (mtp.*)unchanged from the published 4-bit artifact

Size: 19,059,764,358 bytes (19.06 GB), 0.45 GB above the published 4-bit artifact and under the 22 GB listing limit. Byte audit of the saved arrays: 763 arrays, 623 unchanged, 140 changed (the 70 Q8-upgraded Mamba/attention modules, weight + scales); all router gates and all MTP arrays byte-identical to the published 4-bit artifact. Chat template, tokenizer, generation_config.json (temperature 1.0, topp 0.95, dosample true) and the source model card are the vendor's, unchanged; the upstream template at revision a9904d24 is byte-identical to the one shipped here (58933db7…).

config.json carries darkbloom_local_experiment: {profile: "hybrid8", publication_approved: false} and darkbloom_embedded_mtp: {version: 1, architecture: "nemotron_h_attention_moe"}. conversion-receipt.json records the converter (865efefa…), profile rules (0a907d56…), source config hash (a3827a0f…) and per-tensor f32 hashes of every protected tensor.

Files

filesha256
model-00001-of-00004.safetensors9a6f3b9ecbc6036a…
model-00002-of-00004.safetensors545b5e11cd020b30…
model-00003-of-00004.safetensors05858aa57252d4b5…
model-00004-of-00004.safetensorsa02d407d2499028d…
model.safetensors.index.json97251709c2b45e63…
config.jsonc75cadc4b4ddff18…
generation_config.json0ed94823ae737b50…
chat_template.jinja58933db77d3099b4…
tokenizer.json623c34567aebb185…
tokenizer_config.jsonec380f0a6c0c2e29…
conversion-receipt.json3a401b716c7ac8dc…
SOURCE-MODEL-CARD.md5d9d96d2f2b154d7…
LICENSEd7c8a9e5d1896d0a…

Full 64-character digests are in sha256sums.txt in this repository.

Why this build exists

With reasoning disabled (enable_thinking=false), the published 4-bit artifact's greedy continuation for the OpenRouter request "What is the weather like in Boston, MA in fahrenheit?" (one get_current_weather tool, tool_choice auto) is a refusal, every time. The BF16 and Q8 checkpoints call the tool under the same prompt bytes; the refusal is a 4-bit effect. Raising only the Mamba and attention projections to Q8 recovers it.

Validation (loopback Darkbloom provider on an Apple Silicon Mac, 2026-09-12)

Provider binary with the parser fix; prefix cache off; MTP off and on; greedy (temperature 0) unless stated. "Corpora" are Darkbloom's small synthetic tool-selection sets (18 fixed cases, 16 fresh cases, 10 API controls covering tool_choice auto/required/named/none with reasoning on and off, each streaming and non-streaming). They are not a production benchmark.

checkpublished 4-bitthis artifact
exact OpenRouter Boston request, reasoning off, greedyrefusal 3/3 (both MTP modes)tool call 3/3, clean content (both MTP modes)
same request, reasoning ontool calltool call
18-case corpus (discovery / held-out), greedy4/8 · 5/107/8 · 10/10
16 fresh cases (discovery / held-out), greedy8/8 · 4/88/8 · 8/8
API controls, MTP off / on8/10 · 8/1010/10 · 10/10
literal </think> in content (with parser fix)00
exact request sampled at 1.0 / 0.95, 20 per MTP mode5/20 · 10/20 tool calls13/20 · 18/20
MTP on vs offidentical decisionsidentical decisions

Known limitations: the stock-quote style prompt ("current share price of TSLA") still refuses at greedy (1 of 18). Greedy decisions on this model sit on a knife edge: an earlier diagnostic showed the first generated token flipping between a split-prefill Python reference and the provider's full-prefill path on the same token ids, so re-validate on the exact serving path before promotion. Under sampling, some weather prompts refuse that pass at greedy.

License

Released under the NVIDIA Open Model License of the source model; see LICENSE and SOURCE-MODEL-CARD.md. This repository is a private validation artifact of Darkbloom and carries no additional warranty.