CoolFace
Modelpublic

Scicom-intl/Meta-Llama-3.1-70B-Instruct-Malaysian

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes8downloads
Model Card

Qwen2.5-72B-Instruct-Malaysian

SFT LoRA meta-llama/Llama-3.1-70B-Instruct on Scicom-intl/Malaysian-Instructions/commit/288b358a57765a735d588f73e5e6c212c81429bd

  1. 1.Dense LoRA SFT done using DeepSpeed Zero3 HF Trainer.
  2. 2.Multipacking variable length 16384 context length, with global batch size of 32, so global total tokens is 524288.
  3. 3.All linear layers with rank 256 with alpha multiply by 2.0
  4. 4.Liger fused cross entropy.
  5. 5.1e-4 learning rate, 50 warmup, 3 epoch only.

We only upload the best model

<img src="https://raw.githubusercontent.com/Scicom-AI-Enterprise-Organization/small-ablation/refs/heads/main/malaysian-sft/accuracy.png">

Source code

Source code at https://github.com/Scicom-AI-Enterprise-Organization/small-ablation/blob/main/malaysian-sft

Acknowledgement

Special thanks to https://www.scitix.ai/ for H100 Node!