hyjhyj233/fwe-1b-adamw-step100000
021
FWE 1B AdamW — step 100000
Private research checkpoint converted to Hugging Face format.
- Optimizer: AdamW
- Training step: 100000
- Learning rate: 0.003
- Sequence length: 32768
- Model weights, configuration, and tokenizer are included.
- Training optimizer state and intermediate checkpoints are excluded.
