88plug/Kimi-VL-A3B-Thinking-2506-W4A16
034
Kimi-VL-A3B-Thinking-2506-W4A16
W4A16 post-training quantization of moonshotai/Kimi-VL-A3B-Thinking-2506. Language-model Linear layers quantized; vision tower and merger kept BF16.
At a Glance
Quick Start
docker run --gpus device=0 -p 8080:8080 \
vllm/vllm-openai:v0.21.0-cu129-ubuntu2404 vllm serve \
88plug/Kimi-VL-A3B-Thinking-2506-W4A16 \
--trust-remote-code \
--max-model-len 8192 \
--gpu-memory-utilization 0.90Benchmarks
No fabricated numbers. Results will be published to this card when measured.
License
MIT (per base model).
About
**88plug AI Lab** ships compressed-tensors quantizations for native vLLM v0.21.0+ deployment.
This release: Provisional tier — datafree RTN (weight-only rounding, no calibration corpus). A gold AutoRound re-quant is scheduled; 88plug architecture forbids new provisional W4A16 uploads.
Browse all releases → huggingface.co/88plug
