CoolFace
Modelpublic

malaiwah/SmolLM2-135M-QFS-rtn-int4-g64

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
0likes113downloads
Model Card

SmolLM2-135M-QFS-rtn-int4-g64

Uncalibrated round-to-nearest INT4 control. Part of a small QFS stored-weight fidelity campaign, not a broad capability benchmark.

Source: HuggingFaceTB/SmolLM2-135M. Original attribution/card is preserved in ORIGINAL_MODEL_CARD.md. See LICENSE and license-source.json.

All 210 attention/MLP linear matrices use INT4 group size 64. Embeddings, tied output head and norms retain their original BF16 tensor content. GPTQ-v2 packing is decoded by QFS; this does not claim native serving-kernel compatibility.

No calibration or GPTQ optimization ran. quant_method=gptq describes storage compatibility only; the algorithm is RTN.

Changed files: packed model weights and quantization configuration; tokenizer and untouched tensor content are preserved. The conversion receipt describes original conversion outputs before these publication docs were added. scope.json records intervention coverage.

Measurements and calibrated-method comparisons must cite the actual published QFS capture/comparison receipts, not this card alone. QFS Explorer.