CoolFace
Modelpublic

nazdef/gpt2medium-en-it-nanochat-gpt2preln-decay13500-step14250

sourceHugging Facecc-by-sa-4.0updated 3mo agoView on Hugging Face
0likes26downloads
Model Card

GPT2Medium EN/IT NanoChat GPT2PreLN decay-only step_14250

This repository publishes the requested ordinary checkpoint release for step_14250, the current best scalar decayed checkpoint found in the medium GPT2PreLN 13.5k decay family.

This is still not the definitive public-facing family base release.

Released checkpoint summary:

  • —released checkpoint: step_14250.pt
  • —replay continuation run:
  • —20260629_resume-gpt2medium-gpt2preln-k20-wsddecayonly-rerunmissing-lr3p5294e5-anchor20k-final2e5-webwiki-step14200-to14850
  • —original decay-only parent run:
  • —20260628_resume-gpt2medium-gpt2preln-k20-wsddecayonly-lr2e-4-anchor20k-final2e5-webwiki-step13500
  • —parent non-decayed anchor:
  • —stable-recipe-gpt2medium-gpt2preln-k20-wsd-lr2e-4-anchor20k-final2e5-webwiki
  • —original continuation source checkpoint:
  • —step_13500.pt
  • —replay source checkpoint:
  • —step_14200.pt
  • —languages: English + Italian
  • —context window: 2500 tokens
  • —architecture: GPT-2-style decoder with pre-layernorm blocks
  • —architecture config: architecture: gpt2, block_type: gpt2_prelayernorm
  • —training-time parameter count: 337,671,424
  • —published Transformers-export parameter count: 337,639,424
  • —hardware class: single consumer GPU (RTX 4060 Ti 16GB) for training, GPU benchmarked selection

This is a base pretraining checkpoint, not an instruction-tuned chat model.

Why This Checkpoint Exists

The original decay-only tail from step_13500 died on disk-full save failures after step_14200, so the missing cooldown tail was replayed from step_14200 through step_14850.

That replay did not just fill in bookkeeping holes: it produced the current best scalar checkpoint of the entire evaluated decayed 13.5k family.

So this release is not just a provenance artifact. It is the current strongest measured decay-only checkpoint in this family.

Position Inside The Decay Family

The important scoreboard is:

  • —replay step_14250: val_loss_mixed = 4.4419
  • —replay step_14700: val_loss_mixed = 4.4436
  • —replay step_14850: val_loss_mixed = 4.4521
  • —source non-decayed step_13500: val_loss_mixed = 4.4652
  • —old-tail best step_14150: val_loss_mixed = 4.4705
  • —old-tail endpoint step_14200: val_loss_mixed = 4.4967

So the honest read is:

  • —step_14250 is the best scalar checkpoint seen so far in the whole evaluated 13.5k decay family
  • —it beats the source non-decayed step_13500 by about 0.0233 on val_loss_mixed
  • —it beats the best checkpoint from the unreplayed old tail (step_14150) by about 0.0286

Main Metrics For step_14250

  • —val_loss_mixed = 4.4419
  • —val_loss_en = 4.3881
  • —val_loss_it = 3.5969
  • —ppl_mixed = 84.9386

Behavior snapshot:

  • —loop_rate = 0.425
  • —distinct_2 = 0.5617
  • —repeated_4gram_rate = 0.775
  • —language_consistency_en = 0.825
  • —language_consistency_it = 0.950
  • —cloze_en_contains = 0.12
  • —cloze_it_contains = 0.22

Short read:

  • —this is the strongest scalar decay-only checkpoint in the family so far
  • —later-tail checkpoints 14700 and 14850 stay very close, but do not beat it
  • —this is the checkpoint to use when the goal is “best measured decayed release candidate from the 13.5k branch”

Training Data

This model was trained on the bilingual EN/IT web + wiki dataset:

  • —dataset id on disk:
  • —202605141153_fineweb50_wiki50_50en_50it_score100_2500context_5Btokens_tok_20260515_en50it50_webwiki_stratified_500M
  • —context window during training: 2500 tokens
  • —packing length: 2500
  • —mixing strategy: source_balanced
  • —validation ratio: 0.05

Main source groups:

  • —English FineWeb-HQ (epfml/FineWeb-HQ)
  • —Italian FineWeb2-HQ (epfml/FineWeb2-HQ)
  • —English Wiki40B (google/wiki40b)
  • —Italian Wiki40B (google/wiki40b)

How Many Tokens This Checkpoint Saw

Training math:

  • —sequence length: 2500
  • —batch size: 2
  • —grad accumulation: 48
  • —tokens per optimizer step: 239,904

So step_14250 saw approximately:

  • —`3.4186320B` tokens total
  • —`K = 10.1241` tokens per parameter relative to the native training-time parameter count

Included Files

This release bundle includes:

  • —step_14250.pt
  • —step_14250.safetensors
  • —model.safetensors
  • —config.json
  • —tokenizer files
  • —training_config.yaml
  • —run telemetry:
  • —best_validation.json
  • —metrics.jsonl
  • —eval_metrics.jsonl
  • —probe_generations.jsonl
  • —benchmark bundle:
  • —summary.json
  • —comparison.json
  • —comparison.csv
  • —metrics.json
  • —metrics.csv
  • —source_losses.json
  • —report.md
  • —generations.jsonl
  • —generations_comparison.md
  • —cloze_results.jsonl

Quick Start

python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

repo_id = "nazdef/gpt2medium-en-it-nanochat-gpt2preln-decay13500-step14250"

tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)

prompt = "La capitale d'Italia è"
prompt_ids = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
bos = torch.tensor([[tokenizer.bos_token_id]], dtype=prompt_ids["input_ids"].dtype)
input_ids = torch.cat([bos, prompt_ids["input_ids"]], dim=1)
attention_mask = torch.ones_like(input_ids)

outputs = model.generate(
    input_ids=input_ids,
    attention_mask=attention_mask,
    do_sample=True,
    max_new_tokens=64,
    temperature=0.8,
    top_k=50,
    top_p=0.95,
    repetition_penalty=1.1,
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

License

This release is published with `CC-BY-SA-4.0` as the practical downstream posture for the mixed training corpus used here.

The training mix includes:

  • —FineWeb-HQ / FineWeb2-HQ web data
  • —Wiki40B English and Italian slices

Downstream users are responsible for checking whether their use, redistribution, or derivative packaging remains compatible with the obligations of the upstream datasets and their terms.