NexusProjectsAI/Nemotron-3-Nano-30B-A3B-Nexus-Agents-GGUF
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF)
A LoRA fine-tune of nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B (hybrid Mamba/Transformer MoE, ~3B active) specialized for the Nexus Projects agent stack.
Links: the exact training + verification data → Nexus-Agents-ToolCalling · the tool that generated the data, trained, quantized, and evaluated this model → Nexus Training Studio · the app these agents power → Nexus Projects client
- Setup interview — infers industry/platforms/objectives from a free-text idea ("I want to sell lemonade" → Food & Beverage), asks instead of guessing when the input is vague or ambiguous ("a lemonade stand" → asks Food & Beverage or Retail?), and looks libraries up on the internet (pub.dev/GitHub) to pick current ones before finishing.
- Discovery — builds a user-story tree (
As a … I want … so that …). - Task generation — turns setup + stories into concrete, stack-specific tasks with acceptance criteria and verification commands.
How it was trained
Two LoRA stages (rank 16/scale 32, attention + Mamba mixer + always-on shared_experts MLP — never attention-only), on schema-verified synthetic tool-calling conversations carrying the agents' real tool schemas — train == serve. Both corpora are published in the dataset repo:
Results — base vs fine-tuned
Behavioral interview eval (27 end-to-end agent scenarios, greedy decoding)
Each scenario is a full multi-turn setup interview driven against the served GGUF with the agents' real tool schemas. A case passes only if the model covers every required field, never re-asks an answered question, and cleanly finalizes. Full transcripts for every case (base + fine-tuned) are published in the dataset repo's `verification/` folder.
Tool-call accuracy (BFCL-style, 150 held-out calls)
The base model only emitted a parseable tool call 109/150 times; the fine-tune did so 147/150. In short, fine-tuning ~tripled tool-use accuracy on the Nexus tool set.
Quantizations
All quants are imatrix-calibrated.
imatrix.dat (the calibration importance matrix) is included for re-quantizing.
⚠️ Serving requirements
- Disable thinking. Nemotron-3-Nano is a reasoning model; with thinking on it reasons in prose instead of calling tools. Serve with
enable_thinking=false(renders<think></think>). It is also prompt-sensitive — use the agent's real system prompt. - Tool-call format is Nemotron's
<tool_call><function=NAME><parameter=key>value</parameter>…</function></tool_call>— parse that (not JSON) on the client.
License
Inherits the NVIDIA Open Model License from the base model.
