shawnzzzzz/Qwen3-30B-A3B-DAPO-FFN-W4A4-QAT-BF16Master-step-0620
Qwen3-30B-A3B-DAPO-FFN-W4A4-QAT-BF16Master-step-0620
This repository contains training checkpoint step 620, converted from a distributed training checkpoint into standard Hugging Face safetensors. It is part of the Qwen3-30B-A3B W4A4-QAT vs BF16 Checkpoints series.
Checkpoint metadata
- Architecture:
Qwen3MoeForCausalLM - Model type:
qwen3_moe - Base model:
Qwen/Qwen3-30B-A3B-Base - Training trajectory: FFN-only W4A4 quantization-aware training
- Public trajectory label:
W4A4-QAT - Source checkpoint:
global_step_620 - Tensor storage: BF16
- Matched counterpart: shawnzzzzz/Qwen3-30B-A3B-DAPO-BF16-step-0620
- Matched source step:
620 - Absolute step difference:
0
Important quantization note
W4A4-QAT describes how this checkpoint was trained, not its on-disk format. This repository contains clean BF16 master weights and has no quantization_config; it is not a packed W4A4 model. The source training checkpoint did not persist activation-observer scales. Clean-BF16 analysis is supported, while exact W4A4 continuation requires recalibration.
Intended use
These checkpoints are research artifacts for comparing approximately step-matched W4A4-QAT and BF16 training trajectories. They have not been evaluated here as general-purpose production models.
Validation
The export was checked for:
- required Hugging Face model and tokenizer metadata;
- readable safetensors headers and complete shard index;
- exact index-to-shard key consistency;
- BF16 tensor dtype throughout;
- exact key and tensor-shape match against the native Qwen3-MoE architecture;
- full source-artifact content comparison against the validated export;
- Hugging Face path, byte-size, LFS SHA256, and metadata-download integrity.
SHA256SUMS covers every published file in this repository except the checksum manifest itself. Internal execution provenance is intentionally omitted from this public release.
