specklabs/Speck2-140M-GGUF
Speck2-140M GGUF
llama.cpp-compatible GGUF builds of specklabs/Speck2-140M, pinned to source revision 1201df613d9ee9d50909189f52e45c6fcefa3c01.
Usage
llama-completion -hf specklabs/Speck2-140M-GGUF:Q4_K_M -p "The meaning of life is" -n 64The source Speck architecture and llama.cpp's LFM2 runtime implement the same alternating attention/short-convolution operators. Conversion folds the 640-to-768 input and 768-to-640 output adapters into the embeddings, zero-pads the 384-wide convolution channels to 768, and left-pads 3-tap causal kernels to 5 taps. These transformations preserve the model function apart from normal floating-point and quantization rounding.
The GGUF graph stores 180,160,768 parameters because the source's tied 640-wide embedding and two adapters become separate 768-wide input and output matrices. This compatibility transform does not add layers or model capacity.
The conversion was built with llama.cpp revision 2e88c49c90f0add8796f633fea8c3d65b975f295. Exact checksums and conversion provenance are in `conversion.json`.
