CoolFace
Modelpublic

OpenMOSS-Team/SmolLM-1.7B-base-mla-topk2-rank384

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes38downloads
Model Card

SmolLM-1.7B-base-mla-topk2-rank384

This repository contains the attention-only incremental weights for an MLAfication conversion of HuggingFaceTB/SmolLM-1.7B. It is not a standalone full-model checkpoint.

Variant

FieldValue
Base modelHuggingFaceTB/SmolLM-1.7B
Base revisiond7449ff7241c863f3e8accc475155f0f97afa011
MLA latent rank384
RoPE dimensions per KV head2
Training checkpointstep 8000
Training variantstage1+2-distill-QKV
Delta tensors168
Delta size0.49 GiB
SHA-2562a2e1812d02e35cfddbcd915b515dffe8226e1ec9454d3a59f507c8c16e77f09

The training run froze every parameter outside attention and also froze each attention output projection. The uploaded delta therefore contains exactly the parameters matching "attn" in name and "o_proj" not in name; embeddings, MLP, normalization outside attention, output projections, and LM head are omitted. Frozen tensors from the historical full checkpoint were verified exactly against the pinned base weights after casting the base tensors to the checkpoint dtype.

Loading

Use the MLAfication implementation to instantiate the MLAfication architecture from the pinned base model, then load model.safetensors with strict=False:

python
from safetensors.torch import load_file

delta = load_file("model.safetensors")
load_result = mla_model.load_state_dict(delta, strict=False)

The architecture must be patched before loading the delta. Loading this repository directly with vanilla AutoModelForCausalLM.from_pretrained(...) is not supported. config.json contains the layer-wise RoPE indices and sanitized MLAfication metadata.

Files

  • —model.safetensors: attention-only delta weights
  • —config.json: base architecture, RoPE indices, and MLAfication metadata
  • —manifest.json: provenance, byte size, tensor count, and checksum

License

The weights follow the Apache 2.0 license of the SmolLM base model.