CoolFace
Modelpublic

specklabs/Speck2-140M-GGUF

sourceHugging Facemitupdated 12d agoView on Hugging Face
1likes241downloads
Model Card

[image]

Speck2-140M GGUF

llama.cpp-compatible GGUF builds of specklabs/Speck2-140M, pinned to source revision 1201df613d9ee9d50909189f52e45c6fcefa3c01.

FileQuantizationSize
Speck2-140M-BF16.ggufBF16361.2 MB
Speck2-140M-Q4_K_M.ggufQ4KM112.9 MB
Speck2-140M-Q5_K_M.ggufQ5KM130.3 MB
Speck2-140M-Q8_0.ggufQ8_0192.3 MB

Usage

bash
llama-completion -hf specklabs/Speck2-140M-GGUF:Q4_K_M -p "The meaning of life is" -n 64

The source Speck architecture and llama.cpp's LFM2 runtime implement the same alternating attention/short-convolution operators. Conversion folds the 640-to-768 input and 768-to-640 output adapters into the embeddings, zero-pads the 384-wide convolution channels to 768, and left-pads 3-tap causal kernels to 5 taps. These transformations preserve the model function apart from normal floating-point and quantization rounding.

The GGUF graph stores 180,160,768 parameters because the source's tied 640-wide embedding and two adapters become separate 768-wide input and output matrices. This compatibility transform does not add layers or model capacity.

The conversion was built with llama.cpp revision 2e88c49c90f0add8796f633fea8c3d65b975f295. Exact checksums and conversion provenance are in `conversion.json`.