ml-ryanlee/looped-16x2-d640-1e18
027
looped-32L-d640-1e18
32-layer compute-optimal checkpoint for Sparse Layers are Critical to Scaling Looped Language Models (arXiv:2605.09165).
This width is the architecture's measured minimum on a ten-rung isoFLOP sweep (d256-d1408) at 1e18 FLOPs.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained(
"ml-ryanlee/looped-32L-d640-1e18", trust_remote_code=True)
tok = AutoTokenizer.from_pretrained("gpt2")Pass max_length=1024 when evaluating — the RoPE buffer is sized to the training context.
