CoolFace
Modelpublic

Scicom-intl/Qwen3-30B-A3B-Instruct-2507-Malaysian

sourceHugging Faceupdated 9mo agoView on Hugging Face
2likes12downloads
Model Card

Qwen3-30B-A3B-Instruct-2507-Malaysian

SFT LoRA Qwen/Qwen3-30B-A3B-Instruct-2507 on Scicom-intl/Malaysian-Instructions/commit/288b358a57765a735d588f73e5e6c212c81429bd

  1. 1.MoE LoRA SFT done using FSDP2 Fused MoE.
  2. 2.Multipacking variable length 16384 context length, with global batch size of 32, so global total tokens is 524288.
  3. 3.All linear layers with experts, rank 256 with alpha multiply by 2.0 <sup> + </sup>.
  4. 4.Liger fused cross entropy.
  5. 5.1e-4 learning rate, 50 warmup, 3 epoch only.

<sup> + </sup> with the rank of each equal to the total rank divided by the number of active experts, https://thinkingmachines.ai/blog/lora/

We only upload the best model

<img src="https://raw.githubusercontent.com/Scicom-AI-Enterprise-Organization/small-ablation/refs/heads/main/malaysian-sft/accuracy.png">

Source code

Source code at https://github.com/Scicom-AI-Enterprise-Organization/small-ablation/blob/main/malaysian-sft

Acknowledgement

Special thanks to https://www.scitix.ai/ for H100 Node!