malaiwah/SmolLM2-135M-QFS-rtn-int4-g64
SmolLM2-135M-QFS-rtn-int4-g64
Uncalibrated round-to-nearest INT4 control. Part of a small QFS stored-weight fidelity campaign, not a broad capability benchmark.
Source: HuggingFaceTB/SmolLM2-135M. Original attribution/card is preserved in ORIGINAL_MODEL_CARD.md. See LICENSE and license-source.json.
All 210 attention/MLP linear matrices use INT4 group size 64. Embeddings, tied output head and norms retain their original BF16 tensor content. GPTQ-v2 packing is decoded by QFS; this does not claim native serving-kernel compatibility.
No calibration or GPTQ optimization ran. quant_method=gptq describes storage compatibility only; the algorithm is RTN.
Changed files: packed model weights and quantization configuration; tokenizer and untouched tensor content are preserved. The conversion receipt describes original conversion outputs before these publication docs were added. scope.json records intervention coverage.
Measurements and calibrated-method comparisons must cite the actual published QFS capture/comparison receipts, not this card alone. QFS Explorer.
