Spangler3000/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-hybrid8-mtp
NVIDIA Nemotron 3.5 Lightning 30B-A3B — MLX hybrid8 (Mamba/attention Q8, experts Q4) with embedded MTP
Status: private validation candidate. Not the published Darkbloom artifact. Not qualified for production. This checkpoint exists to fix a reasoning-off tool-calling regression observed on the published 4-bit listing. Serve it only with a Darkbloom provider that includes the Nemotron stray-</think> parser fix (private d-inference PR #2, commit A); without that fix this artifact emits a bare </think> before its tool call in reasoning-off mode.
What it is
A mixed-precision MLX conversion of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 at source revision a9904d24bcc1d289a1950fa9d2b978c47cf903b9, with the embedded multi-token-prediction (MTP) head, using the structural hybrid8 rule:
Size: 19,059,764,358 bytes (19.06 GB), 0.45 GB above the published 4-bit artifact and under the 22 GB listing limit. Byte audit of the saved arrays: 763 arrays, 623 unchanged, 140 changed (the 70 Q8-upgraded Mamba/attention modules, weight + scales); all router gates and all MTP arrays byte-identical to the published 4-bit artifact. Chat template, tokenizer, generation_config.json (temperature 1.0, topp 0.95, dosample true) and the source model card are the vendor's, unchanged; the upstream template at revision a9904d24 is byte-identical to the one shipped here (58933db7…).
config.json carries darkbloom_local_experiment: {profile: "hybrid8", publication_approved: false} and darkbloom_embedded_mtp: {version: 1, architecture: "nemotron_h_attention_moe"}. conversion-receipt.json records the converter (865efefa…), profile rules (0a907d56…), source config hash (a3827a0f…) and per-tensor f32 hashes of every protected tensor.
Files
Full 64-character digests are in sha256sums.txt in this repository.
Why this build exists
With reasoning disabled (enable_thinking=false), the published 4-bit artifact's greedy continuation for the OpenRouter request "What is the weather like in Boston, MA in fahrenheit?" (one get_current_weather tool, tool_choice auto) is a refusal, every time. The BF16 and Q8 checkpoints call the tool under the same prompt bytes; the refusal is a 4-bit effect. Raising only the Mamba and attention projections to Q8 recovers it.
Validation (loopback Darkbloom provider on an Apple Silicon Mac, 2026-09-12)
Provider binary with the parser fix; prefix cache off; MTP off and on; greedy (temperature 0) unless stated. "Corpora" are Darkbloom's small synthetic tool-selection sets (18 fixed cases, 16 fresh cases, 10 API controls covering tool_choice auto/required/named/none with reasoning on and off, each streaming and non-streaming). They are not a production benchmark.
Known limitations: the stock-quote style prompt ("current share price of TSLA") still refuses at greedy (1 of 18). Greedy decisions on this model sit on a knife edge: an earlier diagnostic showed the first generated token flipping between a split-prefill Python reference and the provider's full-prefill path on the same token ids, so re-validate on the exact serving path before promotion. Under sampling, some weather prompts refuse that pass at greedy.
License
Released under the NVIDIA Open Model License of the source model; see LICENSE and SOURCE-MODEL-CARD.md. This repository is a private validation artifact of Darkbloom and carries no additional warranty.
