CoolFace
Modelpublic

Raghav-Singhal/epe-1p-smollm-1p7b-100B-20n-2048sl-960gbsz-no_bce

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes28downloads
Model Card

epe-1p-smollm-1p7b-100B-20n-2048sl-960gbsz-no_bce

Converted Hugging Face base checkpoint from the Model Raising EPE pretraining run.

Details

  • —Architecture: LlamaForCausalLM
  • —Base model size: 1.7B
  • —Precision on disk: bfloat16
  • —Source Megatron checkpoint: iter_0050863
  • —Tokenizer: extended SmolLM2 tokenizer with 36 additional special tokens (<assistant> + 35 <charter_X.Y> tokens)
  • —Config vocab size: 49280 padded rows
  • —Tokenizer length: 49188

Variant

This is the 1p EPE variant trained without BCE constitution-prediction loss.

Chat Templates

Two named chat templates are provided:

NameUse case
defaultStandard chat format with the plain assistant role
epeUses <assistant> at the start of assistant turns
python
tok.apply_chat_template(messages, chat_template="default")
tok.apply_chat_template(messages, chat_template="epe")

Always use the bundled tokenizer; the original SmolLM2 tokenizer has only 49152 tokens and will not cover the EPE special tokens.