CoolFace
Modelpublic

TurbulenceDeterministe/Carnice-9b-W8A16-AWQ

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
3likes54downloads
Model Card

Carnice-9b W8A16 AWQ :

8-bit symmetric AWQ quantization of kai-os/Carnice-9b, optimized for Ampere GPUs (RTX 30-series) with vLLM.

How it works :

kai-os/Carnice-9b is a fine-tune of Qwen/Qwen3.5-9B that drops the visual components and uses the Qwen3_5ForCausalLM architecture. This architecture is not natively supported by vLLM.

To work around this, this quantized checkpoint re-wraps the weights back into the Qwen3_5ForConditionalGeneration architecture (matching the original Qwen/Qwen3.5-9B config), so vLLM can load it with --language-model-only to serve text-only inference.

Quantization details:

  • —Method: AWQ (Activation-aware Weight Quantization) via llm-compressor
  • —Bits: 8 (per-channel, symmetric)
  • —Activations: FP16
  • —Ignored layers: linear_attn, lm_head, mtp

Inference Performance :

tested with VLLM bench with a random dataset

HardwareKernelAvg prompt throughput (tokens/s)Avg generation throughput (tokens/s)
One 3090Marlin1993.57221.51
Dual 3090 in 8x8xConch Triton2228.33247.59

Link to the model on Localmaxxing : Carnice-9b W8A16 AWQ

Usage :

VLLM :

On one GPU : ~~~python vllm serve TurbulenceDeterministe/Caranice-9b-W8A16-AWQ --max-model-len auto --reasoning-parser qwen3 --language-model-only #To only load text parameters --tensor-parallel-size 1 ~~~

On multiple GPU : (you need to install the Conch triton Kernel) ~~~python pip install conch-triton-kernels vllm serve TurbulenceDeterministe/Caranice-9b-W8A16-AWQ --max-model-len auto --reasoning-parser qwen3 --language-model-only #To only load text parameters --tensor-parallel-size 2 ~~~