vrfai/Qwen3.6-27B-FP8
Qwen3.6-27B-FP8
FP8 (W8A8) quantized version of Qwen/Qwen3.6-27B by vrfai using llm-compressor.
Also available: vrfai/Qwen3.6-27B-NVFP4 — more aggressive quantization for Blackwell GPUs only.
FP8 Quantization Details
What's Quantized / What's Not
Same selective strategy as the NVFP4 variant — sensitive components are preserved in BF16:
Quantization Config (llm-compressor)
# recipe.yaml
QuantizationModifier:
targets: [Linear]
scheme: FP8
# static W8A8, per-tensor symmetric
ignore:
- lm_head
- re:model\.visual\.blocks\.\d+\..*
- model.visual.merger.linear_fc1
- model.visual.merger.linear_fc2
- re:model\.language_model\.layers\.\d+\.linear_attn\..*Quick Start (vLLM)
vllm serve vrfai/Qwen3.6-27B-FP8 \
--max-model-len 8192 \
--gpu-memory-utilization 0.9 \
--dtype auto \
--trust-remote-code \
--tensor-parallel-size 2Single GPU (≥ 24 GB VRAM, SM 89+):
vllm serve vrfai/Qwen3.6-27B-FP8 \
--max-model-len 8192 \
--gpu-memory-utilization 0.92 \
--dtype auto \
--trust-remote-codeQuantization Script
The recipes and scripts used to quantize this model can be found in the following repository:
Python (Transformers)
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "vrfai/Qwen3.6-27B-FP8"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
messages = [{"role": "user", "content": "Explain quantization in one paragraph."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))OpenAI-compatible API
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="vrfai/Qwen3.6-27B-FP8",
messages=[{"role": "user", "content": "Hello!"}],
temperature=0.7,
max_tokens=512,
)
print(response.choices[0].message.content)NVFP4 vs FP8 Comparison
Tested Environment
Best Practices
Thinking mode:
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
chat_template_kwargs={"enable_thinking": True},
)Credits
- Original model: Qwen Team (Alibaba Group)
- FP8 quantization: vrfai
- Quantization framework: vllm-project/llm-compressor
Below is the original model card from [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B):
<img width="400px" src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.6/logo.png">

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.
Qwen3.6 Highlights
- Agentic Coding: the model now handles frontend workflows and repository-level reasoning with greater fluency and precision.
- Thinking Preservation: reasoning context from historical messages is retained, streamlining iterative development.

For more details, please refer to our blog post Qwen3.6-27B.
Model Overview
- Type: Causal Language Model with Vision Encoder
- Number of Parameters: 27B
- Context Length: 262,144 natively and extensible up to 1,010,000 tokens
Citation
@misc{qwen3.6-27b,
title = {{Qwen3.6-27B}: Flagship-Level Coding in a {27B} Dense Model},
author = {{Qwen Team}},
month = {April},
year = {2026},
url = {https://qwen.ai/blog?id=qwen3.6-27b}
}