CoolFace
Modelpublic

albertge/llada-8b-dllm-memory-tokens-mix60k-recon-w005

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
0likes575downloads
Model Card

llada-8b-dllm-memory-tokens-mix60k-recon-w005

Matched learned memory-token baseline for the dLLM Registers project. This is a memory/compression baseline, not a register-token checkpoint.

This model is described in the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models.

  • —Base model: GSAI-ML/LLaDA-8B-Base
  • —Memory channel: four continuous front-position slots, matching the R4 inference-time channel capacity
  • —Training data: canonical mix60k (OpenMathInstruct-2 + OpenCodeInstruct, 60K examples; source SHA-256 58115f6be4bad1c635f3cdc28a57ee83305d38bb82a10a240030cdd9c32fc7f0)
  • —Chunking: C=128, at most eight chunks, four diffusion-loss passes per trace
  • —Memory objective: reconstruct the preceding completed chunk from the four memory slots, auxiliary weight 0.05
  • —Gradient routing: the next-chunk task loss is detached at the preceding memory writer; reconstruction loss trains the writer
  • —Prompt dropout: 0.3 Bernoulli per trace
  • —Optimization: LR 2e-5, batch size 1, eight MI355X GPUs, seed 42
  • —Matched GSM8K gate (200 examples): carry 106/200 (53.0%), reset 86/200 (43.0%)
  • —Date uploaded: 2026-09-03

The saved config records d1_detach_primary_register_bridge=true, d1_aux_recon_loss=true, d1_aux_recon_weight=0.05, and num_registers=4.

How to load

python
from transformers import AutoModel, AutoTokenizer

repo = "albertge/llada-8b-dllm-memory-tokens-mix60k-recon-w005"
model = AutoModel.from_pretrained(repo, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)

This diffusion checkpoint requires the project evaluator for chunked carry inference; a standard autoregressive generation pipeline is not sufficient. See `eval/eval.py`. Key flags are --num_registers 4 --front_registers --channel_mode registers --tail_length 0 --use_mask_token_for_registers --use_register_carryover.

Repository

Training and evaluation code: https://github.com/SprocketLab/dllm-registers

Citation

bibtex
@misc{ge2026register,
  title  = {Register Tokens for Bounded-State Reasoning in Diffusion Language Models},
  author = {Albert Ge and Chandan Singh and Yufan Zhuang and Xiaodong Liu and Jianfeng Gao and Frederic Sala},
  year   = {2026},
  note   = {Preprint},
  url    = {https://huggingface.co/papers/2609.16372}
}