CoolFace
Modelpublic

88plug/Kimi-VL-A3B-Thinking-2506-W4A16

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes34downloads
Model Card

Kimi-VL-A3B-Thinking-2506-W4A16

W4A16 post-training quantization of moonshotai/Kimi-VL-A3B-Thinking-2506. Language-model Linear layers quantized; vision tower and merger kept BF16.

At a Glance

PropertyValue
Base modelmoonshotai/Kimi-VL-A3B-Thinking-2506
Release tierProvisional (datafree RTN — re-quant scheduled)
Quant methoddatafree RTN W4A16 (AutoRound blocked — KNOWN-FAILURES)
FLAC statusNot measured (T+7d milestone)
Quant formatcompressed-tensors (Marlin G32)
Disk size~18 GB

Quick Start

bash
docker run --gpus device=0 -p 8080:8080 \
  vllm/vllm-openai:v0.21.0-cu129-ubuntu2404 vllm serve \
  88plug/Kimi-VL-A3B-Thinking-2506-W4A16 \
  --trust-remote-code \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.90

Benchmarks

MetricStatus
Throughput (tok/s)In progress — T+7d milestone
MMLU delta vs BF16In progress — T+7d milestone
RULER@128kIn progress — T+30d milestone

No fabricated numbers. Results will be published to this card when measured.

License

MIT (per base model).

About

**88plug AI Lab** ships compressed-tensors quantizations for native vLLM v0.21.0+ deployment.

This release: Provisional tier — datafree RTN (weight-only rounding, no calibration corpus). A gold AutoRound re-quant is scheduled; 88plug architecture forbids new provisional W4A16 uploads.

Browse all releases → huggingface.co/88plug