CoolFace
Modelpublic

OpenMOSS-Team/SmolLM-1.7B-base-mla-topk2-rank640

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes44downloads
Model Card

SmolLM-1.7B-base-mla-topk2-rank640

This repository contains the attention-only incremental weights for an MLAfication conversion of HuggingFaceTB/SmolLM-1.7B. It is not a standalone full-model checkpoint.

Variant

FieldValue
Base modelHuggingFaceTB/SmolLM-1.7B
Base revisiond7449ff7241c863f3e8accc475155f0f97afa011
MLA latent rank640
RoPE dimensions per KV head2
Training checkpointstep 5600
Training variantstage1+2-distill-QKV
Delta tensors168
Delta size0.56 GiB
SHA-256aba8fbaf9ef079a5ca7695c305346c7f1c764dc79435153b6f10352c6594780d

The training run froze every parameter outside attention and also froze each attention output projection. The uploaded delta therefore contains exactly the parameters matching "attn" in name and "o_proj" not in name; embeddings, MLP, normalization outside attention, output projections, and LM head are omitted. Frozen tensors from the historical full checkpoint were verified exactly against the pinned base weights after casting the base tensors to the checkpoint dtype.

Loading

Use the MLAfication implementation to instantiate the MLAfication architecture from the pinned base model, then load model.safetensors with strict=False:

python
from safetensors.torch import load_file

delta = load_file("model.safetensors")
load_result = mla_model.load_state_dict(delta, strict=False)

The architecture must be patched before loading the delta. Loading this repository directly with vanilla AutoModelForCausalLM.from_pretrained(...) is not supported. config.json contains the layer-wise RoPE indices and sanitized MLAfication metadata.

Files

  • —model.safetensors: attention-only delta weights
  • —config.json: base architecture, RoPE indices, and MLAfication metadata
  • —manifest.json: provenance, byte size, tensor count, and checksum

License

The weights follow the Apache 2.0 license of the SmolLM base model.