CoolFace
Modelpublic

rjz123/colar-gsmmath-warmstart-l1b

sourceHugging Facellama3.2updated 1mo agoView on Hugging Face
0likes
Model Card

CoLaR — Math (GSM8K + MATH, warm-start), Llama-3.2-1B

A CoLaR (Compressed Latent Reasoning) checkpoint for Math (GSM8K + Hendrycks MATH), fine-tuned from unsloth/Llama-3.2-1B-Instruct. The model reasons in compressed continuous latent embeddings rather than explicit chain-of-thought tokens.

Model

  • —Base model: unsloth/Llama-3.2-1B-Instruct
  • —Framework: CoLaR — Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains (arXiv:2505.16552).
  • —Domain: Math (GSM8K + Hendrycks MATH)
  • —Warm-start: colar-gsm
  • —Files: colar_hardmath_warmstart_llama1b.ckpt, hparams.yaml

Training procedure

Trained with the CoLaR supervised fine-tuning (SFT) recipe: the frozen base LLM is adapted with q/v LoRA (rank 128, alpha 32) plus a trainable Latent Head (3-layer MLP) and an embedding-compression module. The objective is next-token cross-entropy on the answer plus an `embed_modeling_loss` (MSE) that reconstructs the compressed reasoning-step embeddings (each latent token summarizes about compression_factor chain-of-thought tokens). Reinforcement learning (GRPO) is disabled — this is an SFT-only checkpoint.

  • —This checkpoint: CoLaR SFT, compression factor 4, 12 epochs, warm-started from colar-gsm; SFT only.

Datasets

  • —GSM8K-Aug
  • —Hendrycks MATH

How to load

This is a PyTorch-Lightning checkpoint (weights under the top-level key state_dict) that fits the CoLaR scaffold — it is not directly AutoModel-loadable. Load the base model, splice this state_dict in with strict=False, and use the CoLaR runtime settings:

COLAR_EMB_STD=0.018   COLAR_COMPRESS=<compression_factor>   sep_token=###
TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1

See the CoLaR repository (github.com/xiaomi-research/colar) for the exact loader. The shared GSM8K warm-start ancestor is the official AlbertTan/CoLaR release.

Research artifact for latent-reasoning study (small 1–1.5B model).