EchoLabs33/qwen2.5-7b-instruct-hxq
Qwen2.5-7B-Instruct-HXQ
Qwen2.5-7B-Instruct compressed with HXQ (HelixCode vector quantization). Available as both HuggingFace safetensors (viahelix-substrate) and native GGUF (viallama.cppHXQ fork).
GGUF Runtime Benchmark (RTX 3090 Ti)
Benchmarked against standard GGUF K-quants on RTX 3090 Ti, full GPU offload (-ngl 99), using the `hxq-affine-type` branch at commit 580e9a2.
Decode Speed (tg128, 3 runs)
Perplexity (WikiText-2, 50 chunks, ctx=512)
Prefill (pp512, 3 runs)
Summary: HXQAF6 has the lowest perplexity of all four formats and decodes 15.7% faster than Q6K while being smaller (5.56 vs 5.82 GB). It trades ~10.5% decode speed vs Q4KM for better quality preservation at 6.25 bpw.
This pattern matches the 3B coder results, where HXQ also had the best PPL and fastest decode vs Q6_K.
Reproducibility
All claims are within-run comparisons using the same dataset, llama.cpp commit, and hardware. Do not compare these PPL numbers with numbers from other runs using different model variants, dataset files, or build configurations.
Receipt with SHA256 artifact hashes, exact commands, and dataset provenance: hxq_runtime_3090ti_qwen7b_instruct_20260509
Install and Run
Option 1: Native GGUF (llama.cpp)
# Build llama.cpp with HXQ support
git clone -b hxq-affine-type https://github.com/echo313unfolding/llama.cpp.git
cd llama.cpp && mkdir build && cd build
cmake .. -DGGML_CUDA=ON && make -j$(nproc) llama-cli
# Run
./bin/llama-cli -m qwen2.5-7b-instruct-hxq-affine6.gguf \
-ngl 99 -p "Explain the theory of relativity in simple terms:" -n 128Option 2: HuggingFace (Python)
pip install "helix-substrate[hf]"import helix_substrate # registers the HXQ quantizer with HuggingFace
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("EchoLabs33/qwen2.5-7b-instruct-hxq")
tokenizer = AutoTokenizer.from_pretrained("EchoLabs33/qwen2.5-7b-instruct-hxq")
inputs = tokenizer("Explain the theory of relativity in simple terms:", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))Safetensors Benchmark
Note: The safetensors PPL (7.388) and GGUF PPL (7.982) use different evaluation configurations (ctx=2048/stride=512 vs ctx=512/50 chunks). They are not directly comparable.
Good to Know
- GPU and CPU supported — runs on any CUDA GPU or CPU via standard PyTorch. Native GGUF runs via llama.cpp.
- Fine-tunable via LoRA — compressed weights remain frozen, but LoRA adapters attach to each
HelixLinearlayer viaHelixLinearSTE. Seehelix-substratefor training infrastructure. - Requires `helix-substrate` for safetensors path — the quantizer is not built into transformers.
- Requires llama.cpp HXQ fork for GGUF path — standard llama.cpp does not have HXQ type support yet.
- Tied embeddings —
lm_headsharesembed_tokens, stored at full precision.
What is HXQ?
HXQ is a weight compression codec based on vector quantization with per-group affine correction:
- Each weight matrix is replaced by a 256-entry codebook + uint8 index matrix + per-group affine scale/offset
- The compressed form is the executable —
codebook[indices] * scale + offsetduring matmul, no decompression step - Works on any
nn.Linearregardless of architecture (Transformer, Mamba, MLP) - No calibration data required — codebooks are fit from the weights alone via k-means
- 6.25 bits per weight in the GGUF affine-6 format
Companion Models
Same codec, multiple architectures:
Citation
@software{hxq_2026,
title={HXQ: Vector Quantization with Per-Group Affine Correction for Neural Network Weight Compression},
author={Echo Labs},
year={2026},
url={https://github.com/echo313unfolding/helix-substrate}
}License
Apache 2.0 (inherited from Qwen/Qwen2.5-7B-Instruct).
