nazdef/gpt2medium-en-it-nanochat-gpt2preln-decay13500-step14250
GPT2Medium EN/IT NanoChat GPT2PreLN decay-only step_14250
This repository publishes the requested ordinary checkpoint release for step_14250, the current best scalar decayed checkpoint found in the medium GPT2PreLN 13.5k decay family.
This is still not the definitive public-facing family base release.
Released checkpoint summary:
- released checkpoint:
step_14250.pt - replay continuation run:
20260629_resume-gpt2medium-gpt2preln-k20-wsddecayonly-rerunmissing-lr3p5294e5-anchor20k-final2e5-webwiki-step14200-to14850- original decay-only parent run:
20260628_resume-gpt2medium-gpt2preln-k20-wsddecayonly-lr2e-4-anchor20k-final2e5-webwiki-step13500- parent non-decayed anchor:
stable-recipe-gpt2medium-gpt2preln-k20-wsd-lr2e-4-anchor20k-final2e5-webwiki- original continuation source checkpoint:
step_13500.pt- replay source checkpoint:
step_14200.pt- languages: English + Italian
- context window:
2500tokens - architecture: GPT-2-style decoder with pre-layernorm blocks
- architecture config:
architecture: gpt2,block_type: gpt2_prelayernorm - training-time parameter count:
337,671,424 - published Transformers-export parameter count:
337,639,424 - hardware class: single consumer GPU (
RTX 4060 Ti 16GB) for training, GPU benchmarked selection
This is a base pretraining checkpoint, not an instruction-tuned chat model.
Why This Checkpoint Exists
The original decay-only tail from step_13500 died on disk-full save failures after step_14200, so the missing cooldown tail was replayed from step_14200 through step_14850.
That replay did not just fill in bookkeeping holes: it produced the current best scalar checkpoint of the entire evaluated decayed 13.5k family.
So this release is not just a provenance artifact. It is the current strongest measured decay-only checkpoint in this family.
Position Inside The Decay Family
The important scoreboard is:
- replay
step_14250:val_loss_mixed = 4.4419 - replay
step_14700:val_loss_mixed = 4.4436 - replay
step_14850:val_loss_mixed = 4.4521 - source non-decayed
step_13500:val_loss_mixed = 4.4652 - old-tail best
step_14150:val_loss_mixed = 4.4705 - old-tail endpoint
step_14200:val_loss_mixed = 4.4967
So the honest read is:
step_14250is the best scalar checkpoint seen so far in the whole evaluated13.5kdecay family- it beats the source non-decayed
step_13500by about0.0233onval_loss_mixed - it beats the best checkpoint from the unreplayed old tail (
step_14150) by about0.0286
Main Metrics For step_14250
val_loss_mixed = 4.4419val_loss_en = 4.3881val_loss_it = 3.5969ppl_mixed = 84.9386
Behavior snapshot:
loop_rate = 0.425distinct_2 = 0.5617repeated_4gram_rate = 0.775language_consistency_en = 0.825language_consistency_it = 0.950cloze_en_contains = 0.12cloze_it_contains = 0.22
Short read:
- this is the strongest scalar decay-only checkpoint in the family so far
- later-tail checkpoints
14700and14850stay very close, but do not beat it - this is the checkpoint to use when the goal is “best measured decayed release candidate from the
13.5kbranch”
Training Data
This model was trained on the bilingual EN/IT web + wiki dataset:
- dataset id on disk:
202605141153_fineweb50_wiki50_50en_50it_score100_2500context_5Btokens_tok_20260515_en50it50_webwiki_stratified_500M- context window during training:
2500tokens - packing length:
2500 - mixing strategy:
source_balanced - validation ratio:
0.05
Main source groups:
- English FineWeb-HQ (
epfml/FineWeb-HQ) - Italian FineWeb2-HQ (
epfml/FineWeb2-HQ) - English Wiki40B (
google/wiki40b) - Italian Wiki40B (
google/wiki40b)
How Many Tokens This Checkpoint Saw
Training math:
- sequence length:
2500 - batch size:
2 - grad accumulation:
48 - tokens per optimizer step:
239,904
So step_14250 saw approximately:
- `3.4186320B` tokens total
- `K = 10.1241` tokens per parameter relative to the native training-time parameter count
Included Files
This release bundle includes:
step_14250.ptstep_14250.safetensorsmodel.safetensorsconfig.json- tokenizer files
training_config.yaml- run telemetry:
best_validation.jsonmetrics.jsonleval_metrics.jsonlprobe_generations.jsonl- benchmark bundle:
summary.jsoncomparison.jsoncomparison.csvmetrics.jsonmetrics.csvsource_losses.jsonreport.mdgenerations.jsonlgenerations_comparison.mdcloze_results.jsonl
Quick Start
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
repo_id = "nazdef/gpt2medium-en-it-nanochat-gpt2preln-decay13500-step14250"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)
prompt = "La capitale d'Italia è"
prompt_ids = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
bos = torch.tensor([[tokenizer.bos_token_id]], dtype=prompt_ids["input_ids"].dtype)
input_ids = torch.cat([bos, prompt_ids["input_ids"]], dim=1)
attention_mask = torch.ones_like(input_ids)
outputs = model.generate(
input_ids=input_ids,
attention_mask=attention_mask,
do_sample=True,
max_new_tokens=64,
temperature=0.8,
top_k=50,
top_p=0.95,
repetition_penalty=1.1,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))License
This release is published with `CC-BY-SA-4.0` as the practical downstream posture for the mixed training corpus used here.
The training mix includes:
- FineWeb-HQ / FineWeb2-HQ web data
- Wiki40B English and Italian slices
Downstream users are responsible for checking whether their use, redistribution, or derivative packaging remains compatible with the obligations of the upstream datasets and their terms.
