nazdef/gpt2small-en-it-nanochat-gpt2preln-cpt8600-lr5e5-predecay-step18000
GPT2PreLN Small EN/IT CPT Checkpoint step_18000
This repository packages the current best continual-pretraining checkpoint before any decay-only extraction in the small GPT2PreLN EN/IT line.
Concretely, this is:
- source family base:
nazdef/1gpu-llm-small-en-it-base - source checkpoint lineage:
step_8600.pt - continual-pretraining winner:
step_18000.pt - learning-rate regime for the continuation branch:
5e-5 - decay-only status: not done yet
So this repo should be read as:
a public CPT experiment checkpoint derived from 1gpu-llm-small-en-it-base, not yet the final post-decay family-base promotion.What This Is
1gpu-llm-small-en-it-base was the previous public small-family winner, produced by a short decay-only continuation up to step_8600.
This repo starts from that released checkpoint and continues pretraining with a lower learning rate and a new local schedule:
- source checkpoint:
step_8600.pt - continuation mode:
continual_pretraining - continuation LR:
5e-5 - continuation schedule:
- warmup:
500 - stable:
10000 - decay:
1000
The resulting branch was evaluated checkpoint-by-checkpoint on CPU, and step_18000 emerged as the best checkpoint of the branch.
Architecture
- architecture family: GPT-2-style decoder
- block type:
gpt2_prelayernorm - context window:
2500 - vocab size:
32000 - dim:
768 - n_layers:
12 - n_heads:
12 - parameter count:
136,128,000(~136.128M)
This is a base model, not an instruction-tuned chat model.
Provenance
- public source repo:
nazdef/1gpu-llm-small-en-it-base - local source checkpoint:
step_8600.pt - source parent run:
20260622_resume-gpt2small-gpt2preln-k20-wsds800-final2e5-webwiki-step8000-dense50- continual-pretraining run:
20260701_continual-pretraining-gpt2small-step8600-lr5e5-w500-s10000-d1000-final1e5-webwiki- released checkpoint from this branch:
step_18000.pt
Important release caveat:
- this branch has not gone through the final
wsd-decay-onlyextraction stage yet step_18000is therefore the best pre-decay continual-pretraining checkpoint, not the final cooled-down release candidate
Training Data
This branch stayed on the same bilingual EN/IT web+wiki family used by the small base release:
- dataset id:
202605141153_fineweb50_wiki50_50en_50it_score100_2500context_5Btokens_tok_20260515_en50it50_webwiki_stratified_500M- mixing strategy:
source_balanced - context length during training:
2500 - validation ratio:
0.05
Main source groups:
- English FineWeb-HQ (
epfml/FineWeb-HQ) - Italian FineWeb2-HQ (
epfml/FineWeb2-HQ) - English Wiki40B (
google/wiki40b) - Italian Wiki40B (
google/wiki40b)
Token Accounting
Training math:
- sequence length:
2500 - batch size:
6 - grad accumulation:
16 - tokens per optimizer step:
240,000
Approximate token exposure:
- source
step_8600: about 2.064B tokens - extra tokens added by continual pretraining from
8600 -> 18000: about 2.256B tokens - total by
step_18000: about 4.320B tokens
Why step_18000 Was Chosen
The branch was evaluated in four benchmark passes:
12000/12500quick batch- remaining mid-branch batch
- late-branch batch
- final
step_20000
Lower val_loss_mixed is better.
Key checkpoints:
- source
step_8600:4.7964 - branch
step_9500:4.7565 - branch
step_12500:4.8020 - branch
step_18000:4.7358 - branch
step_20000:4.7892
So:
- the lower-LR continual-pretraining branch did beat the source release
- the best point was not the terminal checkpoint
- the branch peak was
step_18000
Direct Comparison vs 1gpu-llm-small-en-it-base
1gpu-llm-small-en-it-base corresponds to the source step_8600 checkpoint.
Main metrics
Reading
Where step_18000 is better:
val_loss_mixed:4.7964 -> 4.7358val_loss_en:4.8075 -> 4.7260ppl_mixed:121.07 -> 113.96loop_rate:0.525 -> 0.450repeated_4gram_rate:0.925 -> 0.875cloze_it_contains:0.20 -> 0.28
Where step_8600 is still better:
val_loss_it:3.6415 vs 3.8244language_consistency_it:0.825 vs 0.775
Short honest read:
- yes, this branch produced a real new winner
- no, it is not a clean sweep across every metric
- the gain is strongest on the mixed benchmark and repetition/loop cleanup
- the main remaining tradeoff is weaker Italian-side scalar consistency than the old public base
Probe Snapshot at step_18000
The capital of Italy is-> expectedRomecorrect_token_rank = 7correct_token_probability = 0.0172perplexity_on_target_sequence = 58.10A small language model should-> expectedbecorrect_token_rank = 1correct_token_probability = 0.5547perplexity_on_target_sequence = 1.80La capitale d'Italia è-> expectedRomacorrect_token_rank = 5correct_token_probability = 0.0356perplexity_on_target_sequence = 28.05Un piccolo modello linguistico dovrebbe-> expectedesserecorrect_token_rank = 1correct_token_probability = 0.3320perplexity_on_target_sequence = 3.01
Practical reading:
- procedural prompts remain top-1
- factual prompts are still fragile and somewhat loopy
- the branch is stronger overall than the source, but not yet “done”
Recommended Decoding
No dedicated decoding sweep has been run yet on step_18000.
For now, this repo ships the same conservative public preset used by nazdef/1gpu-llm-small-en-it-base:
- preset:
balanced do_sample = truetemperature = 0.8top_k = 50top_p = 0.95repetition_penalty = 1.1no_repeat_ngram_size = 0max_new_tokens = 64
This is an inherited safe default, not a claim that step_18000 already has a dedicated decoding-search winner.
Quick Start
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
repo_id = "nazdef/gpt2small-en-it-nanochat-gpt2preln-cpt8600-lr5e5-predecay-step18000"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)
prompt = "A small language model should"
prompt_ids = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
bos = torch.tensor([[tokenizer.bos_token_id]], dtype=prompt_ids["input_ids"].dtype)
input_ids = torch.cat([bos, prompt_ids["input_ids"]], dim=1)
attention_mask = torch.ones_like(input_ids)
outputs = model.generate(
input_ids=input_ids,
attention_mask=attention_mask,
do_sample=True,
max_new_tokens=64,
temperature=0.8,
top_k=50,
top_p=0.95,
repetition_penalty=1.1,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))Files Included
- original
.ptcheckpoint - exported checkpoint-native
.safetensorsweights plus metadata sidecar - standard Transformers
model.safetensors - Transformers
config.json - tokenizer files
- continual-pretraining config
- branch telemetry:
best_validation.jsonmetrics.jsonleval_metrics.jsonlprobe_generations.jsonl- selected eval artifacts for the branch
- branch closeout report
- probe summary for
step_18000 generation_config.jsonrecommended_decoding_params.jsonrelease_note.md
Intended Use
Use this model as:
- a public continual-pretraining checkpoint derived from
1gpu-llm-small-en-it-base - a checkpoint to inspect before the decay-only extraction stage
- a likely anchor candidate for a later
wsd-decay-onlycooldown tail
Do not read this repo as:
- the final replacement for
1gpu-llm-small-en-it-base - the final post-decay public family-base release
That final decision belongs to the next decay-only stage.
