XXXXyu/bitnet_b1_58-3B-vlut-gguf
074
bitnetb158-3B-vlut-gguf
This repository contains state-of-the-art ternary-packed versions of bitnet_b1_58-3B in GGUF format, optimized for efficient on-device inference using the Vec-LUT method.
Key Features
- ๐ฏ SOTA Compression: Achieves BPW (bits per weight) as low as 1.60 through lossless sub-2-bit ternary packing.
- โก SOTA Performance: Delivers superior throughput (4.2x speedup) in parallel inference scenarios via vector lookup table (LUT).
- ๐ Drop-in Ready: Seamless integration with vlut.cpp for immediate deployment on edge devices.
Available Model Variants
Models are named as ggml-model-{PACKING}_{TILE}.gguf:
Selection Guide
- BPW vs. Speed:
I1_Vachieves lower memory usage but may not always outperformI2_Vin speed. - Tiling Trade-off: Tiled variants (tile size > 1) deliver higher throughput but require larger cache capacity.
- Starting Point: Use
I1_V_2orI2_V_4as a starting point.
For detailed tiling parameter analysis, see Evaluation.md and the paper.
Usage
Prerequisites
Install vlut.cpp (these models require vlut.cpp, not vanilla llama.cpp):
git clone https://github.com/Cipherxzc/vlut.cpp.git
cd vlut.cpp
cmake -B build && cmake --build build --config Release -j4Download & Run
# Download the recommended variant, e.g., I2_V_4
hf download <repo_id> \
ggml-model-I2_V_4.gguf --local-dir ./models
# Run parallel inference
./build/bin/llama-batched \
-m ./models/ggml-model-I2_V_4.gguf \
-p "I believe the meaning of life is" \
-np 32 -n 16 -t 1 --temp 0.5 --repeat-penalty 1.5
# Benchmark performance
./build/bin/llama-bench \
-m ./models/ggml-model-I2_V_4.gguf \
-t 1 -p 128 -n 0For comprehensive usage instructions, refer to the vlut.cpp Quick Start Guide.
Citation
If you use these models, please cite our paper:
@article{li2025veclut,
title={Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices},
author={Li, Xiangyu and Yin, Chengyu and Wang, Weijun and Wei, Jianyu and Cao, Ting and Liu, Yunxin},
journal={arXiv preprint arXiv:2512.06443},
year={2025},
url={https://arxiv.org/abs/2512.06443}
}