CoolFace
Modelpublic

Raghav-Singhal/epe-1p-smollm-1p7b-100B-20n-2048sl-960gbsz-no_bce-refl_end_training

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes21downloads
Model Card

epe-1p-smollm-1p7b-100B-20n-2048sl-960gbsz-nobce-reflend_training

Converted Hugging Face base checkpoint from the Model Raising pretraining run.

Details

  • —Architecture: LlamaForCausalLM
  • —Base model size: 1.7B
  • —Precision on disk: bfloat16
  • —Source Megatron checkpoint iteration: 5863
  • —Model kind: epe
  • —Config vocab size: 49280

Tokenizer

Use the bundled tokenizer from this repository.

This EPE checkpoint uses the extended SmolLM2 tokenizer with <assistant> and 35 <charter_X.Y> tokens. Two named chat templates are available:

NameAssistant turn start
default`<im_start>assistant\n`
epe`<im_start><assistant>\n`