CoolFace
Modelpublic

carlopires/Qwopus3.8-27B-Flash-PTQ-groupwise-int-NInfer

sourceHugging Faceapache-2.0updated 22d agoView on Hugging Face
1likes237downloads
Model Card

Qwopus3.8-27B-Flash PTQ groupwise-int for NInfer

This release packages Qwopus3.8-27B-Flash for NInfer using post-training quantization (PTQ) to NInfer's groupwise-int weight storage. It did not undergo quantization-aware training (QAT).

Qwopus3.8-27B-Flash is an agent-oriented "Flash" fine-tune of Qwen/Qwen3.8-27B published by Jackrong as GGUF only. The source checkpoint has no public Safetensors distribution, so this artifact was built by converting the publisher's Qwopus3.8-27B-Flash-MTP-Q5_K_M.gguf to a canonical BF16 HF checkpoint (merging the official base-model Vision tower and frontend resources) and converting to the NInfer container with CPU groupwise-int quantization.

The artifact is intended only for NInfer. It is not a Transformers checkpoint, Safetensors distribution, GGUF file, or generic interchange format. NInfer selects it with --model qwopus3.8-27b (the fork's default).

A faster NVFP4 variant of the same fine-tune is available at carlopires/Qwopus3.8-27B-Flash-PTQ-NVFP4-NInfer; on spot measurements it decodes about 1.63x faster than this groupwise-int artifact with MTP enabled.

Artifact

FieldValue
Filenameqwopus3_8_27b.ninfer
Size18,210,531,328 bytes (16.96 GiB)
SHA-2568110c98efdf6cc6722f2f850a7630aa6c503b69c3fa060d88a505ae11e25e5ee
Container version2
NInfer model IDqwopus3.8-27b
NInfer weights IDgroupwise-int
NInfer target keyqwen3_8_27b
Conversion recipeqwen3_8_27b-v1
Stored objects1,124 (1,118 tensors and 6 frontend resources)

The file contains the Text backbone, Vision tower, MTP draft model, optimized proposal head, tokenizer, chat template, generation configuration, and media processor resources required by NInfer.

Verify the download with:

bash
printf '%s  %s\n' \
  '8110c98efdf6cc6722f2f850a7630aa6c503b69c3fa060d88a505ae11e25e5ee' \
  'qwopus3_8_27b.ninfer' | sha256sum --check

Requirements

  • —64-bit Linux;
  • —an NVIDIA Blackwell GPU with FP4 support (the NInfer runtime targets sm_120a);
  • —CUDA Toolkit 13.1 or newer;
  • —NInfer with qwopus3.8-27b registered, from carlopires/ninfer-rtx5090-mobile, commit `830e26bb`.

The artifact used 15.92 GiB of GPU weight storage during validation. Context capacity requires additional memory. A 24 GiB RTX 5090 Laptop GPU supports useful contexts with INT8 KV; a 32 GiB desktop RTX 5090 provides more context headroom and is NInfer's primary performance target.

Download and run

Build the runtime:

bash
git clone https://github.com/carlopires/ninfer-rtx5090-mobile.git
cd ninfer-rtx5090-mobile
git checkout 830e26bb34504640d0e7fabcd203b1e7a81d60ff

cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Download the artifact:

bash
hf download carlopires/Qwopus3.8-27B-Flash-PTQ-groupwise-int-NInfer \
  qwopus3_8_27b.ninfer \
  --local-dir models

Run the CLI with MTP speculative decoding:

bash
./build/apps/ninfer models/qwopus3_8_27b.ninfer \
  --model qwopus3.8-27b \
  --prompt "Explain quantization-aware training in three sentences." \
  --max-context 4096 --max-new 256 \
  --kv-dtype int8 \
  --spec mtp --draft-tokens 3 --lm-head-draft

--model qwopus3.8-27b is this fork's default and may be omitted. Loading the published Qwen3.8 QUASAR artifact on the same build requires --model qwen3.8-27b; a mismatched selector is rejected with an identity error.

Run the OpenAI/Anthropic-compatible server:

bash
./build/apps/ninfer-serve models/qwopus3_8_27b.ninfer \
  --host 127.0.0.1 --port 18080 \
  --max-concurrency 1 --max-context 4096 \
  --kv-dtype int8 \
  --spec mtp --draft-tokens 3 --lm-head-draft

Increase context and concurrency only within the available VRAM. NInfer fixes active-request capacity at startup.

Supported NInfer features

This is a complete Qwen3.8 artifact rather than a Text-only weight file. The registered route supports:

  • —thinking and non-thinking text generation;
  • —image, multi-image, video, and mixed multimodal messages;
  • —MTP speculative decoding with optimized proposal-head lookup;
  • —BF16 and INT8 group-64 KV cache;
  • —CUDA Graph decode and compatible-prefix reuse;
  • —the NInfer CLI and bounded concurrent server;
  • —OpenAI Responses Core, OpenAI Chat Completions, and Anthropic Messages translation.

NInfer is a single-GPU engine. It does not provide CPU/GPU offload, multi-GPU execution, distributed serving, or large-scale preemptive continuous batching.

How it was built

Qwopus3.8-27B-Flash publishes GGUF weights only, so the conversion ran in two stages.

Stage 1 — GGUF to BF16 HF

Qwopus3.8-27B-Flash-MTP-Q5_K_M.gguf (arch qwen35, 65 blocks = 64 text layers + 1 MTP layer, 16 full-attention + 48 Gated DeltaNet layers) was exported with tools/convert/qwen3_8_27b/gguf_export.py, which applies the inverse transforms of llama.cpp's Qwen3.5 mapping: GDN V-row reorder inversion, out_proj column inverse, conv1d channel inverse, ssm_a = -exp(A_log) recovery, and the llama.cpp norm +1 offset reversal. The result is 1,199 BF16 tensors: 866 Text/MTP tensors plus the 333 model.visual.* tensors from the fixed official Qwen3.8 checkpoint (Qwopus shares the base vision tower).

Stage 2 — NInfer groupwise-int conversion

The standard qwen3_8_27b-v1 recipe converted the 1,199-tensor BF16 checkpoint on CPU (344 s), producing 1,124 stored objects. The groupwise-int weight formats are:

FormatTensors
BF16582
FP3296
I321
Q4G64_F16S183
Q5G64_F16S246
Q6G64_F16S1
W8G32_F16S9

Fidelity and quality caveats

Quantization and evaluation scope

This is a PTQ release, not a QAT release, and it carries a fidelity ceiling: the source was a lossy Q5KM GGUF, dequantized to BF16 and then re-quantized to groupwise-int storage. Quality approximates the publisher's fine-tune; it is not bit-faithful to it. Results published for the upstream fine-tune or other deployments are not measurements of this artifact. Its quality must be assessed separately.

Qwopus publishes no HF Safetensors ground truth, so no numeric comparison against the original fine-tune was possible beyond the GGUF source itself. No broad accuracy or long-context evaluation has been run on this artifact. A higher-fidelity rebuild from the publisher's BF16 GGUF (or the NVFP4 variant) would change the cost/quality balance.

Validation performed

  • —the artifact loads and generates coherent text on the RTX 5090 Laptop through the public Engine route (~39 tok/s decode, INT8 KV, spot check), including MTP speculative decode;
  • —identity enforcement verified in both directions (qwopus3.8-27b accepted, qwen3.8-27b rejected against this artifact);
  • —converter preflight passed over the complete 1,124-object plan.

A successful load and coherent responses do not demonstrate benchmark quality.

Reproduce the conversion

Download the pinned inputs:

bash
hf download Qwen/Qwen3.8-27B \
  --revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 \
  --local-dir models/Qwen3.8-27B

hf download Jackrong/Qwopus3.8-27B-Flash-GGUF \
  Qwopus3.8-27B-Flash-MTP-Q5_K_M.gguf \
  --revision 3c24631f3a14fb86682bc0c60224e63d42dd5bf1 \
  --local-dir models/qwopus

Check out the converter and install its Python dependencies in a Python 3.11 environment:

bash
git clone https://github.com/carlopires/ninfer-rtx5090-mobile.git
cd ninfer-rtx5090-mobile
git checkout 830e26bb34504640d0e7fabcd203b1e7a81d60ff

python3.11 -m venv .venv
.venv/bin/python -m pip install torch numpy safetensors gguf==0.19.0

Export and convert:

bash
# GGUF -> canonical BF16 HF (Text/MTP tensors)
PYTHONPATH=. .venv/bin/python -m tools.convert.qwen3_8_27b.gguf_export \
  --gguf models/qwopus/Qwopus3.8-27B-Flash-MTP-Q5_K_M.gguf \
  --out models/qwopus/hf_bf16

# merge Vision tensors + frontend resources into the full HF directory, then:
OMP_NUM_THREADS=2 PYTHONPATH=. .venv/bin/python -m tools.convert.qwen3_8_27b.convert \
  --model models/qwopus/model_hf \
  --out qwopus3_8_27b.ninfer \
  --device cpu \
  --model-id qwopus3.8-27b

The converter validates the checkpoint configuration, every selected source tensor and dtype, the frontend resources, and the complete 1,124-object plan before opening the output.

Provenance

RoleRepositoryRevision
Base modelQwen/Qwen3.8-27B1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
Fine-tune source (GGUF)Jackrong/Qwopus3.8-27B-Flash-GGUF3c24631f3a14fb86682bc0c60224e63d42dd5bf1
Converter/runtimecarlopires/ninfer-rtx5090-mobile830e26bb34504640d0e7fabcd203b1e7a81d60ff

See qwopus3_8_27b.ninfer.conversion.json for the generated conversion report and artifact-manifest.json for a compact machine-readable summary.

License

This artifact is distributed under the Apache License 2.0, matching the Qwen3.8 base model and the Qwopus3.8-27B-Flash source repository. Users remain responsible for complying with the source licenses and applicable laws. If this artifact is useful, cite the original Qwen3.8 model and Jackrong's Qwopus3.8-27B-Flash work.