specklabs/Speck2-140M-Instruct-GGUF
Speck2-140M-Instruct GGUF
llama.cpp-compatible GGUF builds of specklabs/Speck2-140M-Instruct, pinned to source revision 1d959ef826b1160754a967aed9536aea0a76eedc.
Usage
llama-cli -hf specklabs/Speck2-140M-Instruct-GGUF:Q4_K_M -cnvThe source Speck architecture and llama.cpp's LFM2 runtime implement the same alternating attention/short-convolution operators. Conversion folds the 640-to-768 input and 768-to-640 output adapters into the embeddings, zero-pads the 384-wide convolution channels to 768, and left-pads 3-tap causal kernels to 5 taps. These transformations preserve the model function apart from normal floating-point and quantization rounding.
The GGUF graph stores 180,165,376 parameters because the source's tied 640-wide embedding and two adapters become separate 768-wide input and output matrices. This compatibility transform does not add layers or model capacity.
The conversion was built with llama.cpp revision 2e88c49c90f0add8796f633fea8c3d65b975f295. Exact checksums and conversion provenance are in `conversion.json`.
