huluhuluu/qwen3-4b-instruct-2507-eagle3-sharegpt-lr5e-5-epoch0-step10000
Qwen3-4B-Instruct-2507 EAGLE3 ShareGPT (LR 5e-5)
Online EAGLE3 draft-model training run with SpecForge using a peak learning rate of 5e-5. This archive contains the two available checkpoints, epoch_0_step_5000 and epoch_0_step_10000; each is published as a separate Hub model repository in the companion collection.
This is a speculative-decoding draft model, not a standalone chat model. Pair it with the exact target model family.
Training parameters
Architecture
LlamaForCausalLMEagle3, one decoder layer, hidden size 2560, intermediate size 9728, 32 attention heads, 8 key/value heads, draft vocabulary size 32000, target vocabulary size 151936, bfloat16 weights. The standard run does not set a sliding-window limit.
Checkpoint files
Every checkpoint repository contains model.safetensors, config.json, and training_state.pt. The latter stores optimizer/scheduler state and training arguments for resuming and should only be deserialized in a trusted environment. Prefer model.safetensors for inference.
Usage
Use a checkpoint repository as the SGLang speculative draft path with Qwen/Qwen3-4B-Instruct-2507 and the EAGLE3 speculative-decoding settings supported by your SGLang version. Tree settings should be benchmarked for the serving workload. No evaluation or safety metrics were recorded for this run.
