specklabs/Speck1-140M-Instruct-GGUF
Speck1-140M-Instruct GGUF
llama.cpp-compatible GGUF builds of specklabs/Speck1-140M-Instruct, pinned to source revision 16ad80599d499490b70317770a84a18466719bba.
Usage
llama-cli -hf specklabs/Speck1-140M-Instruct-GGUF:Q4_K_M -cnvThe source Speck architecture and llama.cpp's LFM2 runtime implement the same alternating attention/short-convolution operators. Conversion folds the 640-to-768 input and 768-to-640 output adapters into the embeddings, zero-pads the 384-wide convolution channels to 768, and left-pads 3-tap causal kernels to 5 taps. These transformations preserve the model function apart from normal floating-point and quantization rounding.
The GGUF graph stores 180,165,376 parameters because the source's tied 640-wide embedding and two adapters become separate 768-wide input and output matrices. This compatibility transform does not add layers or model capacity.
The conversion was built with llama.cpp revision 2e88c49c90f0add8796f633fea8c3d65b975f295. Exact checksums and conversion provenance are in `conversion.json`.
