CoolFace
Modelpublic

OpenMOSS-Team/Qwen2.5-7B-base-mla-topk8-rank512

sourceHugging Faceapache-2.0updated 10d agoView on Hugging Face
0likes40downloads
Model Card

Qwen2.5-7B-base-mla-topk8-rank512

This repository contains the attention-only incremental weights for an MLAfication conversion of Qwen/Qwen2.5-7B. It is not a standalone full-model checkpoint.

Variant

FieldValue
Base modelQwen/Qwen2.5-7B
Base revisiond149729398750b98c0af14eb82c78cfe92750796
MLA latent rank512
RoPE dimensions per KV head8
Training checkpointstep 180
Training variantstage1+2-distill-QKV
Delta tensors336
Delta size1.47 GiB
SHA-25609e3c288e055711f3413457e32f0f3397ead889d472b15ddabb625c6f96dc684

The training run froze every parameter outside attention and also froze each attention output projection. The uploaded delta therefore contains exactly the parameters matching "attn" in name and "o_proj" not in name; embeddings, MLP, normalization outside attention, output projections, and LM head are omitted. Frozen tensors from the historical full checkpoint were verified exactly against the pinned base weights after casting the base tensors to the checkpoint dtype.

Loading

Use the MLAfication implementation to instantiate the MLAfication architecture from the pinned base model, then load model.safetensors with strict=False:

python
from safetensors.torch import load_file

delta = load_file("model.safetensors")
load_result = mla_model.load_state_dict(delta, strict=False)

The architecture must be patched before loading the delta. Loading this repository directly with vanilla AutoModelForCausalLM.from_pretrained(...) is not supported. config.json contains the layer-wise RoPE indices and sanitized MLAfication metadata.

Files

  • —model.safetensors: attention-only delta weights
  • —config.json: base architecture, RoPE indices, and MLAfication metadata
  • —manifest.json: provenance, byte size, tensor count, and checksum

License

The weights follow the Apache 2.0 license of the Qwen2.5 base model.