geodesic-research/nemotron-super-120b-cc-mt-replay-only-1b-sr-sft
07
nemotron-super-120b-cc-mt-replay-only-1b-sr-sft
Nemotron 3 Super 120B-A12B reasoning (short-CoT <think> traces) SFT from the constitutional-curriculum midtraining study — the SFT stage applied to the cc-mt-replay-only-1b midtrained base.
Training
- Base: geodesic-research/nemotron-super-120b-cc-mt-replay-only-1b (re-imported to Megatron, fresh optimizer warm-start)
- Data: geodesic-research/persistent-alignment-warm-start-short-reasoning
default(982.5M tokens, 113,387 packed rows @8192) - Recipe: 886 iterations, GBS 128, seq 8192, lr 5e-6 cosine w/ 10% warmup, tokenizer
geodesic-research/nemotron-think-tokenizer, answer-only loss. Final lm loss 0.684. - Topology: TP=1 EP=4 PP=22 ETP=1, 22 nodes / 88 GH200 GPUs on Isambard-AI.
- Coherence: 8/8 clean at T=0.6 topp=0.95, full 8192-token budget — [W&B bq7f4iyj](https://wandb.ai/geodesic/megatronbridgeconversioncoherance_tests/runs/bq7f4iyj)
Usage notes
Chat model — apply the bundled chat template. Recommended sampling: temperature 0.6, top_p 0.95 (Nemotron recommendation; higher temperatures occasionally emit a stray token right before EOS). Loads with native transformers NemotronHForCausalLM; tokenizer PreTrainedTokenizerFast.
