CoolFace
Modelpublic

ml-ryanlee/looped-16x2-d640-1e18

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes27downloads
Model Card

looped-32L-d640-1e18

32-layer compute-optimal checkpoint for Sparse Layers are Critical to Scaling Looped Language Models (arXiv:2605.09165).

architecturelooped
d_model640
effective layers32
compute budget1e18 FLOPs
training steps45,632
parameters (stored)143,627,520
peak LR0.005
batch size16
muP width_ratio2.5 (d_base=256)
final val lossn/a

This width is the architecture's measured minimum on a ten-rung isoFLOP sweep (d256-d1408) at 1e18 FLOPs.

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained(
    "ml-ryanlee/looped-32L-d640-1e18", trust_remote_code=True)
tok = AutoTokenizer.from_pretrained("gpt2")

Pass max_length=1024 when evaluating — the RoPE buffer is sized to the training context.