CoolFace
Modelpublic

XReyRobert/Nex-N2-mini-GPTQ-Pro

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
2likes30downloads
Model Card

[image]

Nex-N2-mini GPTQ-Pro

<p style="margin: 10px 0 6px 0; color: #475569; font-size: 14px; line-height: 1.5;"> These models are built and maintained on rented GPU compute. If you want to show some appreciation, a follow on X or a coffee helps keep the releases coming. </p>

<div style="display: flex; flex-wrap: wrap; align-items: center; gap: 10px; margin: 8px 0 18px 0;"> <a href="https://x.com/xreyrobert" target="blank" rel="noopener noreferrer" class="follow-link" style="display: inline-flex; align-items: center; justify-content: center; gap: 8px; min-height: 36px; box-sizing: border-box; padding: 9px 14px; border: 1px solid #111827; border-radius: 999px; background: #111827; color: #ffffff; text-decoration: none; font-weight: 700; font-size: 14px; line-height: 1;"> <svg viewBox="0 0 24 24" fill="currentColor" aria-hidden="true" style="width: 16px; height: 16px; flex: 0 0 auto;"><path d="M18.244 2.25h3.308l-7.227 8.26 8.502 11.24H16.17l-5.214-6.817L4.99 21.75H1.68l7.73-8.835L1.254 2.25H8.08l4.713 6.231zm-1.161 17.52h1.833L7.084 4.126H5.117z"></path></svg> Follow @xreyrobert </a> <a href="https://buymeacoffee.com/xrrxrr" target="blank" rel="noopener noreferrer" class="support-link" style="display: inline-flex; align-items: center; justify-content: center; gap: 8px; min-height: 36px; box-sizing: border-box; padding: 9px 14px; border: 1px solid #f59e0b; border-radius: 999px; background: #fbbf24; color: #111827; text-decoration: none; font-weight: 700; font-size: 14px; line-height: 1;"> Buy me a coffee </a> </div>

This is a GPTQ-Pro 4-bit quantization of `nex-agi/Nex-N2-mini`.

It is a deployment artifact, not a new fine-tune. The goal is to make the Nex-N2-mini MoE checkpoint easier to test in GPTQ-compatible local serving stacks while keeping the model card honest about the validation status.

The source checkpoint includes vision/visual tensors. This artifact preserves those tensors, but the validated publication story here is text and coding-agent serving. Vision behavior has not yet been validated for the quantized artifact.

Source And Credits

Source model:

Quantization tooling and reference recipe:

Artifact Summary

FieldValue
Source modelnex-agi/Nex-N2-mini
ArchitectureQwen3_5MoeForConditionalGeneration
Model typeqwen3_5_moe
Tensor files5
Safetensors size19.23 GiB
Indexed tensors124576
Quantized qweight tensors30970
mtp.* tensors in indexfalse
vision/visual tensors in indextrue
Index metadata size matches shardstrue

The source index/logs showed no mtp.* tensors. This artifact therefore normalizes text_config.mtp_num_hidden_layers to 0 and records the change under artifact_notes.mtp.

Quantization Recipe

SettingValue
MethodGPTQ-Pro / GPTQModel
Quantizergptqmodel:6.1.0-dev
Bits4
Group size128
Symmetric quantizationtrue
Desc actfalse
True sequentialtrue
Calibration datasetWikiText
Calibration samples256
Calibration sequence length2048
MSE2.0
Damp percent0.05
Damp auto increment0.01
FOEM alpha0.25
FOEM beta0.2
FOEM devicecuda:0
MoE routingExpertsRoutingBypass
MoE bypass batch size320
Dense VRAM strategyexclusive
MoE VRAM strategybalanced
Pack implementationcpu

Fallback smoothing was enabled for difficult groups with threshold 0.5%.

Intended Serving Shape

This checkpoint is intended for advanced users testing text-only GPTQ serving for Qwen3.6-style MoE models.

A starting vLLM shape for text-only testing:

bash
vllm serve XReyRobert/Nex-N2-mini-GPTQ-Pro \
  --served-model-name nex-n2-mini-gptq-pro \
  --language-model-only \
  --dtype float16 \
  --quantization gptq_marlin \
  --tensor-parallel-size 1 \
  --max-model-len 262144 \
  --max-num-seqs 1 \
  --kv-cache-dtype fp8_e5m2 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --enable-prefix-caching \
  --gpu-memory-utilization 0.95 \
  --trust-remote-code

Serving context for the published Smoke24/vLLM measurements:

The Smoke24/vLLM numbers were collected on an internal llm-residency vLLM deployment. The custom image recipe is not published yet, so this card does not present that image as a public reproduction target. The stable serving knobs captured from the run are listed for context.

FieldValue
Nomad job profilevllm-nex-n2-mini-262k
Served model namenex-n2-mini-gptq-pro-ctx262k
Critical flags--dtype float16, --quantization gptq_marlin, --kv-cache-dtype fp8_e5m2, --reasoning-parser qwen3, --tool-call-parser qwen3_coder, --max-model-len 262144, --max-num-batched-tokens 2096
Benchmark contextSmoke24 quality rows used max_model_len=131072 for apples-to-apples comparison; the image above reflects the validated 262k residency serving profile.

Treat this as a starting point. Loader compatibility depends on vLLM, Transformers, GPTQModel, GPTQ-Marlin, and Qwen3.6 MoE support.

The RTX 3090 image above reflects separate 262k-context serving validation.

Public vLLM Reproducibility

This artifact has a public reproducibility path on the unmodified upstream vLLM OpenAI image:

  • —image: docker.io/vllm/vllm-openai:nightly-7a1eb8ac2ec4ea69338c51dc7afd4b15010abfa8
  • —vLLM version observed in validation: 0.20.1rc1.dev16+g7a1eb8ac2
  • —GPU class: single RTX 3090 24 GB / Ampere
  • —--enforce-eager was not used
  • —no local sleep/wake patch or localhost/*sleepwake* image is required for the validation below

Validated serving shapes:

  • —--max-model-len 131072 validated with --gpu-memory-utilization 0.94
  • —--max-model-len 262144 validated with --gpu-memory-utilization 0.96
  • —--language-model-only, --dtype float16, --quantization gptq_marlin
  • —--kv-cache-dtype fp8_e5m2, --enable-prefix-caching, --max-num-seqs 1
  • —--max-num-batched-tokens 2096, --max-cudagraph-capture-size 32
  • —--reasoning-parser qwen3, --tool-call-parser qwen3_coder

The 262k profile is tight on 24 GB GPUs; the lower 0.94 memory target used for 131k was short on KV cache, while 0.96 passed.

Validation And Benchmarks

Completed artifact checks:

  • —Local shard index inspection completed before upload.
  • —Remote file list verified after upload.
  • —Remote model.safetensors.index.json verified after upload.
  • —Index metadata total size matches the local safetensor shards.
  • —The remote artifact contains the expected five safetensor shards.

Terminal-Bench 2.0 Smoke24 result and associated vLLM serving measurements. This Smoke24 run used max_model_len=131072 for apples-to-apples comparison with the other local models in this publication batch:

RunScoreSuccess rateWall-timeOutput tokensObserved decodeLLM API time
nex-n2-mini-gptq-pro14/2458.3%314.6m1670.6k140.8 tok/s197.4m

Smoke24 is a fixed 24-task Terminal-Bench 2.0 comparison corpus, not a full Terminal-Bench leaderboard run. In this harness, Nex-N2-mini GPTQ-Pro tied the Qwen3.6 27B GPTQ reference on solved tasks but used more wall time and far more output tokens. That makes it a useful candidate for further serving and generation-control tuning, not an efficiency leader in this specific test.

Task list and harness shape:

  • —`benchmarks/terminal-bench-2.0/smoke24_task_list_20260616.md`

MTP And Vision Status

  • —mtp.* tensors are not present in this artifact.
  • —text_config.mtp_num_hidden_layers was normalized to 0.
  • —Do not enable MTP speculative decoding for this artifact.
  • —Vision/visual tensors are present, but multimodal serving has not been validated for this quantized artifact.

Limitations

  • —Experimental quantization.
  • —Terminal-Bench Smoke24 is a small local comparison corpus, not a full benchmark submission.
  • —Nex-N2-mini was verbose and reasoning-heavy in the Smoke24 harness; generation controls may need further tuning.
  • —MTP speculative decoding is not supported by this artifact.
  • —Vision tensors are preserved, but vision behavior has not been validated.
  • —Loader behavior may vary across vLLM, Transformers, GPTQModel, and GPTQ-Marlin versions.

Files

Key files:

  • —model.safetensors.index.json
  • —model-00001-of-00005.safetensors through model-00005-of-00005.safetensors
  • —config.json
  • —quantize_config.json
  • —processor_config.json
  • —tokenizer.json
  • —UPLOAD_MANIFEST.json

UPLOAD_MANIFEST.json records the upload guardrail checks and artifact inspection summary.

References

Individual Project Notice

This repository is an individual research project. It is not affiliated with, sponsored by, or endorsed by any employer or organization.