Scicom-intl/Qwen3-30B-A3B-Instruct-2507-Malaysian-DoRA
15
Qwen3-30B-A3B-Instruct-2507-Malaysian-DoRA
SFT DoRA Qwen/Qwen3-30B-A3B-Instruct-2507 on Scicom-intl/Malaysian-Instructions/commit/288b358a57765a735d588f73e5e6c212c81429bd
- MoE DoRA SFT done using FSDP2 Fused MoE.
- Multipacking variable length 16384 context length, with global batch size of 32, so global total tokens is 524288.
- All linear layers with experts, rank 256 with alpha multiply by 2.0 <sup> + </sup>.
- Liger fused cross entropy.
- 1e-4 learning rate, 50 warmup, 3 epoch only.
<sup> + </sup> with the rank of each equal to the total rank divided by the number of active experts, https://thinkingmachines.ai/blog/lora/
We only upload the best model
<img src="https://raw.githubusercontent.com/Scicom-AI-Enterprise-Organization/small-ablation/refs/heads/main/malaysian-sft/accuracy.png">
Source code
Source code at https://github.com/Scicom-AI-Enterprise-Organization/small-ablation/blob/main/malaysian-sft
Acknowledgement
Special thanks to https://www.scitix.ai/ for H100 Node!
