CoolFace
Modelpublic

pico-lm/pico-decoder-medium

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes14kdownloads
Model Card

Pico Decoder Medium

pico-decoder-medium is a 181M parameter model in the pico-decoder suite, balancing scale and analyzability. Built with `pico-train` and instrumented with `pico-analyze`, it enables detailed studies of layer-wise learning behavior during language model pretraining.

NOTE: The pico-decoder-medium-1 branch contains the full commit history for the training run.

๐Ÿ”ง Model Details

FieldValue
ArchitectureDecoder-only transformer (LLaMA-style)
Parameters181M
Layers12
Hidden Size768
Feed Forward Size3072
Attention Heads12
Key/Value Heads4

๐Ÿ“š Training

  • โ€”Dataset: `pretokenized-dolma`
  • โ€”Training steps: 200,000
  • โ€”Batch size: 1024
  • โ€”Sequence length: 2048
  • โ€”Optimizer: AdamW
  • โ€”Learning rate schedule: Linear decay with warmup
  • โ€”Compute: 16 A100-SXM4-80GB GPUs

๐Ÿ“ˆ Evaluation and Analysis

This model supports fine-grained analysis using pico-analyze. This tool enables researchers to understand how learning unfolds over training, even at very small scales.

We also evaluate perplexity of the model on the pico-paloma-tinsy dataset.

๐Ÿ“„ Citation

bibtex
@software{pico2025,
    author = {Diehl Martinez, Richard},
    title = {Pico: A Lightweight Framework for Studying Language Model Learning Dynamics},
    year = {2025},
    url = {https://github.com/pico-lm}
}
pico-lm/pico-decoder-medium ยท CoolFace