CoolFace
Modelpublic

nazdef/gpt2small-en-it-nanochat-gpt2preln-cpt8600-lr5e5-predecay-step18000

sourceHugging Facecc-by-sa-4.0updated 3mo agoView on Hugging Face
0likes20downloads
Model Card

GPT2PreLN Small EN/IT CPT Checkpoint step_18000

This repository packages the current best continual-pretraining checkpoint before any decay-only extraction in the small GPT2PreLN EN/IT line.

Concretely, this is:

  • —source family base: nazdef/1gpu-llm-small-en-it-base
  • —source checkpoint lineage: step_8600.pt
  • —continual-pretraining winner: step_18000.pt
  • —learning-rate regime for the continuation branch: 5e-5
  • —decay-only status: not done yet

So this repo should be read as:

a public CPT experiment checkpoint derived from 1gpu-llm-small-en-it-base, not yet the final post-decay family-base promotion.

What This Is

1gpu-llm-small-en-it-base was the previous public small-family winner, produced by a short decay-only continuation up to step_8600.

This repo starts from that released checkpoint and continues pretraining with a lower learning rate and a new local schedule:

  • —source checkpoint: step_8600.pt
  • —continuation mode: continual_pretraining
  • —continuation LR: 5e-5
  • —continuation schedule:
  • —warmup: 500
  • —stable: 10000
  • —decay: 1000

The resulting branch was evaluated checkpoint-by-checkpoint on CPU, and step_18000 emerged as the best checkpoint of the branch.

Architecture

  • —architecture family: GPT-2-style decoder
  • —block type: gpt2_prelayernorm
  • —context window: 2500
  • —vocab size: 32000
  • —dim: 768
  • —n_layers: 12
  • —n_heads: 12
  • —parameter count: 136,128,000 (~136.128M)

This is a base model, not an instruction-tuned chat model.

Provenance

  • —public source repo: nazdef/1gpu-llm-small-en-it-base
  • —local source checkpoint: step_8600.pt
  • —source parent run:
  • —20260622_resume-gpt2small-gpt2preln-k20-wsds800-final2e5-webwiki-step8000-dense50
  • —continual-pretraining run:
  • —20260701_continual-pretraining-gpt2small-step8600-lr5e5-w500-s10000-d1000-final1e5-webwiki
  • —released checkpoint from this branch:
  • —step_18000.pt

Important release caveat:

  • —this branch has not gone through the final wsd-decay-only extraction stage yet
  • —step_18000 is therefore the best pre-decay continual-pretraining checkpoint, not the final cooled-down release candidate

Training Data

This branch stayed on the same bilingual EN/IT web+wiki family used by the small base release:

  • —dataset id:
  • —202605141153_fineweb50_wiki50_50en_50it_score100_2500context_5Btokens_tok_20260515_en50it50_webwiki_stratified_500M
  • —mixing strategy: source_balanced
  • —context length during training: 2500
  • —validation ratio: 0.05

Main source groups:

  • —English FineWeb-HQ (epfml/FineWeb-HQ)
  • —Italian FineWeb2-HQ (epfml/FineWeb2-HQ)
  • —English Wiki40B (google/wiki40b)
  • —Italian Wiki40B (google/wiki40b)

Token Accounting

Training math:

  • —sequence length: 2500
  • —batch size: 6
  • —grad accumulation: 16
  • —tokens per optimizer step: 240,000

Approximate token exposure:

  • —source step_8600: about 2.064B tokens
  • —extra tokens added by continual pretraining from 8600 -> 18000: about 2.256B tokens
  • —total by step_18000: about 4.320B tokens

Why step_18000 Was Chosen

The branch was evaluated in four benchmark passes:

  • —12000/12500 quick batch
  • —remaining mid-branch batch
  • —late-branch batch
  • —final step_20000

Lower val_loss_mixed is better.

Key checkpoints:

  • —source step_8600: 4.7964
  • —branch step_9500: 4.7565
  • —branch step_12500: 4.8020
  • —branch step_18000: 4.7358
  • —branch step_20000: 4.7892

So:

  • —the lower-LR continual-pretraining branch did beat the source release
  • —the best point was not the terminal checkpoint
  • —the branch peak was step_18000

Direct Comparison vs 1gpu-llm-small-en-it-base

1gpu-llm-small-en-it-base corresponds to the source step_8600 checkpoint.

Main metrics

checkpointval_loss_enval_loss_itval_loss_mixedppl_mixedcloze_it_containslanguage_consistency_itrepeated_4gram_rateloop_rate
step_86004.80753.64154.7964121.06940.200.8250.9250.525
step_180004.72603.82444.7358113.95710.280.7750.8750.450

Reading

Where step_18000 is better:

  • —val_loss_mixed: 4.7964 -> 4.7358
  • —val_loss_en: 4.8075 -> 4.7260
  • —ppl_mixed: 121.07 -> 113.96
  • —loop_rate: 0.525 -> 0.450
  • —repeated_4gram_rate: 0.925 -> 0.875
  • —cloze_it_contains: 0.20 -> 0.28

Where step_8600 is still better:

  • —val_loss_it: 3.6415 vs 3.8244
  • —language_consistency_it: 0.825 vs 0.775

Short honest read:

  • —yes, this branch produced a real new winner
  • —no, it is not a clean sweep across every metric
  • —the gain is strongest on the mixed benchmark and repetition/loop cleanup
  • —the main remaining tradeoff is weaker Italian-side scalar consistency than the old public base

Probe Snapshot at step_18000

  • —The capital of Italy is -> expected Rome
  • —correct_token_rank = 7
  • —correct_token_probability = 0.0172
  • —perplexity_on_target_sequence = 58.10
  • —A small language model should -> expected be
  • —correct_token_rank = 1
  • —correct_token_probability = 0.5547
  • —perplexity_on_target_sequence = 1.80
  • —La capitale d'Italia è -> expected Roma
  • —correct_token_rank = 5
  • —correct_token_probability = 0.0356
  • —perplexity_on_target_sequence = 28.05
  • —Un piccolo modello linguistico dovrebbe -> expected essere
  • —correct_token_rank = 1
  • —correct_token_probability = 0.3320
  • —perplexity_on_target_sequence = 3.01

Practical reading:

  • —procedural prompts remain top-1
  • —factual prompts are still fragile and somewhat loopy
  • —the branch is stronger overall than the source, but not yet “done”

Recommended Decoding

No dedicated decoding sweep has been run yet on step_18000.

For now, this repo ships the same conservative public preset used by nazdef/1gpu-llm-small-en-it-base:

  • —preset: balanced
  • —do_sample = true
  • —temperature = 0.8
  • —top_k = 50
  • —top_p = 0.95
  • —repetition_penalty = 1.1
  • —no_repeat_ngram_size = 0
  • —max_new_tokens = 64

This is an inherited safe default, not a claim that step_18000 already has a dedicated decoding-search winner.

Quick Start

python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

repo_id = "nazdef/gpt2small-en-it-nanochat-gpt2preln-cpt8600-lr5e5-predecay-step18000"

tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)

prompt = "A small language model should"
prompt_ids = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
bos = torch.tensor([[tokenizer.bos_token_id]], dtype=prompt_ids["input_ids"].dtype)
input_ids = torch.cat([bos, prompt_ids["input_ids"]], dim=1)
attention_mask = torch.ones_like(input_ids)

outputs = model.generate(
    input_ids=input_ids,
    attention_mask=attention_mask,
    do_sample=True,
    max_new_tokens=64,
    temperature=0.8,
    top_k=50,
    top_p=0.95,
    repetition_penalty=1.1,
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Files Included

  • —original .pt checkpoint
  • —exported checkpoint-native .safetensors weights plus metadata sidecar
  • —standard Transformers model.safetensors
  • —Transformers config.json
  • —tokenizer files
  • —continual-pretraining config
  • —branch telemetry:
  • —best_validation.json
  • —metrics.jsonl
  • —eval_metrics.jsonl
  • —probe_generations.jsonl
  • —selected eval artifacts for the branch
  • —branch closeout report
  • —probe summary for step_18000
  • —generation_config.json
  • —recommended_decoding_params.json
  • —release_note.md

Intended Use

Use this model as:

  • —a public continual-pretraining checkpoint derived from 1gpu-llm-small-en-it-base
  • —a checkpoint to inspect before the decay-only extraction stage
  • —a likely anchor candidate for a later wsd-decay-only cooldown tail

Do not read this repo as:

  • —the final replacement for 1gpu-llm-small-en-it-base
  • —the final post-decay public family-base release

That final decision belongs to the next decay-only stage.