XReyRobert/Nex-N2-mini-GPTQ-Pro
Nex-N2-mini GPTQ-Pro
<p style="margin: 10px 0 6px 0; color: #475569; font-size: 14px; line-height: 1.5;"> These models are built and maintained on rented GPU compute. If you want to show some appreciation, a follow on X or a coffee helps keep the releases coming. </p>
<div style="display: flex; flex-wrap: wrap; align-items: center; gap: 10px; margin: 8px 0 18px 0;"> <a href="https://x.com/xreyrobert" target="blank" rel="noopener noreferrer" class="follow-link" style="display: inline-flex; align-items: center; justify-content: center; gap: 8px; min-height: 36px; box-sizing: border-box; padding: 9px 14px; border: 1px solid #111827; border-radius: 999px; background: #111827; color: #ffffff; text-decoration: none; font-weight: 700; font-size: 14px; line-height: 1;"> <svg viewBox="0 0 24 24" fill="currentColor" aria-hidden="true" style="width: 16px; height: 16px; flex: 0 0 auto;"><path d="M18.244 2.25h3.308l-7.227 8.26 8.502 11.24H16.17l-5.214-6.817L4.99 21.75H1.68l7.73-8.835L1.254 2.25H8.08l4.713 6.231zm-1.161 17.52h1.833L7.084 4.126H5.117z"></path></svg> Follow @xreyrobert </a> <a href="https://buymeacoffee.com/xrrxrr" target="blank" rel="noopener noreferrer" class="support-link" style="display: inline-flex; align-items: center; justify-content: center; gap: 8px; min-height: 36px; box-sizing: border-box; padding: 9px 14px; border: 1px solid #f59e0b; border-radius: 999px; background: #fbbf24; color: #111827; text-decoration: none; font-weight: 700; font-size: 14px; line-height: 1;"> Buy me a coffee </a> </div>
This is a GPTQ-Pro 4-bit quantization of `nex-agi/Nex-N2-mini`.
It is a deployment artifact, not a new fine-tune. The goal is to make the Nex-N2-mini MoE checkpoint easier to test in GPTQ-compatible local serving stacks while keeping the model card honest about the validation status.
The source checkpoint includes vision/visual tensors. This artifact preserves those tensors, but the validated publication story here is text and coding-agent serving. Vision behavior has not yet been validated for the quantized artifact.
Source And Credits
Source model:
Quantization tooling and reference recipe:
Artifact Summary
The source index/logs showed no mtp.* tensors. This artifact therefore normalizes text_config.mtp_num_hidden_layers to 0 and records the change under artifact_notes.mtp.
Quantization Recipe
Fallback smoothing was enabled for difficult groups with threshold 0.5%.
Intended Serving Shape
This checkpoint is intended for advanced users testing text-only GPTQ serving for Qwen3.6-style MoE models.
A starting vLLM shape for text-only testing:
vllm serve XReyRobert/Nex-N2-mini-GPTQ-Pro \
--served-model-name nex-n2-mini-gptq-pro \
--language-model-only \
--dtype float16 \
--quantization gptq_marlin \
--tensor-parallel-size 1 \
--max-model-len 262144 \
--max-num-seqs 1 \
--kv-cache-dtype fp8_e5m2 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--enable-prefix-caching \
--gpu-memory-utilization 0.95 \
--trust-remote-codeServing context for the published Smoke24/vLLM measurements:
The Smoke24/vLLM numbers were collected on an internal llm-residency vLLM deployment. The custom image recipe is not published yet, so this card does not present that image as a public reproduction target. The stable serving knobs captured from the run are listed for context.
Treat this as a starting point. Loader compatibility depends on vLLM, Transformers, GPTQModel, GPTQ-Marlin, and Qwen3.6 MoE support.
The RTX 3090 image above reflects separate 262k-context serving validation.
Public vLLM Reproducibility
This artifact has a public reproducibility path on the unmodified upstream vLLM OpenAI image:
- image:
docker.io/vllm/vllm-openai:nightly-7a1eb8ac2ec4ea69338c51dc7afd4b15010abfa8 - vLLM version observed in validation:
0.20.1rc1.dev16+g7a1eb8ac2 - GPU class: single RTX 3090 24 GB / Ampere
--enforce-eagerwas not used- no local sleep/wake patch or
localhost/*sleepwake*image is required for the validation below
Validated serving shapes:
--max-model-len 131072validated with--gpu-memory-utilization 0.94--max-model-len 262144validated with--gpu-memory-utilization 0.96--language-model-only,--dtype float16,--quantization gptq_marlin--kv-cache-dtype fp8_e5m2,--enable-prefix-caching,--max-num-seqs 1--max-num-batched-tokens 2096,--max-cudagraph-capture-size 32--reasoning-parser qwen3,--tool-call-parser qwen3_coder
The 262k profile is tight on 24 GB GPUs; the lower 0.94 memory target used for 131k was short on KV cache, while 0.96 passed.
Validation And Benchmarks
Completed artifact checks:
- Local shard index inspection completed before upload.
- Remote file list verified after upload.
- Remote
model.safetensors.index.jsonverified after upload. - Index metadata total size matches the local safetensor shards.
- The remote artifact contains the expected five safetensor shards.
Terminal-Bench 2.0 Smoke24 result and associated vLLM serving measurements. This Smoke24 run used max_model_len=131072 for apples-to-apples comparison with the other local models in this publication batch:
Smoke24 is a fixed 24-task Terminal-Bench 2.0 comparison corpus, not a full Terminal-Bench leaderboard run. In this harness, Nex-N2-mini GPTQ-Pro tied the Qwen3.6 27B GPTQ reference on solved tasks but used more wall time and far more output tokens. That makes it a useful candidate for further serving and generation-control tuning, not an efficiency leader in this specific test.
Task list and harness shape:
- `benchmarks/terminal-bench-2.0/smoke24_task_list_20260616.md`
MTP And Vision Status
mtp.*tensors are not present in this artifact.text_config.mtp_num_hidden_layerswas normalized to0.- Do not enable MTP speculative decoding for this artifact.
- Vision/visual tensors are present, but multimodal serving has not been validated for this quantized artifact.
Limitations
- Experimental quantization.
- Terminal-Bench Smoke24 is a small local comparison corpus, not a full benchmark submission.
- Nex-N2-mini was verbose and reasoning-heavy in the Smoke24 harness; generation controls may need further tuning.
- MTP speculative decoding is not supported by this artifact.
- Vision tensors are preserved, but vision behavior has not been validated.
- Loader behavior may vary across vLLM, Transformers, GPTQModel, and GPTQ-Marlin versions.
Files
Key files:
model.safetensors.index.jsonmodel-00001-of-00005.safetensorsthroughmodel-00005-of-00005.safetensorsconfig.jsonquantize_config.jsonprocessor_config.jsontokenizer.jsonUPLOAD_MANIFEST.json
UPLOAD_MANIFEST.json records the upload guardrail checks and artifact inspection summary.
References
- Source model: `nex-agi/Nex-N2-mini`
- GPTQ-Pro tooling: `groxaxo/GPTQ-Pro`
- Reference recipe: `groxaxo/Qwen3.6-27B-GPTQ-Pro-4bit`
- Terminal-Bench: `laude-institute/terminal-bench`
Individual Project Notice
This repository is an individual research project. It is not affiliated with, sponsored by, or endorsed by any employer or organization.
