qingfengyuhuoda/vllm-sliced-qwen2.5-14b-v12cuda
SliceGPT Qwen2.5-14B-Instruct v12cuda
This is the v12cuda SliceGPT checkpoint derived from Qwen/Qwen2.5-14B-Instruct. It is a structurally compressed model, not a drop-in dense Qwen2 model: the hidden width varies by layer and the residual paths use stored SliceGPT shortcut matrices.
Contents
model.safetensors: fp16 converted checkpoint (26.3 GB).config.json: includes the complete per-layerslicing_configandauto_map.configuration_slicegpt_qwen2.pyandmodeling_slicegpt_qwen2.py: portable Transformers reference implementation used bytrust_remote_code=True.- tokenizer and generation configuration copied from the Qwen2.5 base model.
Transformers loading
Use trust_remote_code=True; loading it as ordinary Qwen2ForCausalLM is incorrect because normal Qwen2 layers do not implement the SliceGPT shortcut or original-width RMS normalization.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_dir = "qingfengyuhuoda/vllm-sliced-qwen2.5-14b-v12cuda"
tokenizer = AutoTokenizer.from_pretrained(
model_dir, trust_remote_code=True, local_files_only=True
)
model = AutoModelForCausalLM.from_pretrained(
model_dir,
trust_remote_code=True,
local_files_only=True,
low_cpu_mem_usage=True,
torch_dtype="auto",
)The bundled Transformers implementation is a correctness/reference runtime. For efficient serving, apply the included downstream vLLM/vLLM-Ascend patch: slicegpt_qwen2_14b_vllm_ascend_patch_for_downstream.
Measured result and limitation
The v12cuda checkpoint was compressed with exact CUDA eigh, PCA orientation, and 128 calibration samples. Under the same-GPU serial CUDA throughput test (GPU5, TP=1, eager, 200 requests × 512 input + 128 output tokens, gpu_memory_utilization=0.60), it obtained 3105.11 total tok/s and 621.02 output tok/s, versus dense 2994.97 / 598.99 (+3.68%).
Quality is not dense-equivalent: remote full WikiText-2 PPL is 8.479382 versus 5.693846 for dense (+48.92%). Use this release for research and deployment integration experiments, not as a quality-neutral replacement for the dense model.
License and attribution
The checkpoint and tokenizer are derived from Qwen2.5-14B-Instruct. Comply with the base model's license and usage terms when using or redistributing this artifact. SliceGPT conversion code is from the associated experiment project.
