CoolFace
Modelpublic

tuandunghcmut/Qwen3-8B-Private

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes28downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

<img src="https://raw.githubusercontent.com/axolotl-ai-cloud/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl" width="200" height="32"/> <details><summary>See axolotl config</summary>

axolotl version: 0.13.0.dev0

yaml
base_model: Qwen/Qwen3-8B
# Automatically upload checkpoint and final model to HF
# hub_model_id: username/custom_model_name
hub_model_id: tuandunghcmut/Qwen3-8B-Private


plugins:
  - axolotl.integrations.cut_cross_entropy.CutCrossEntropyPlugin
strict: false



chat_template: qwen3
datasets:
  # - path: trendmicro-ailab/Primus-Seed
  #   type: chat_template
    # split: train[:20%]
    # split: train
    # field_messages: conversations
    # message_property_mappings:
    #   role: from
    #   content: value
  - path: trendmicro-ailab/Primus-Reasoning
    type: chat_template
    # split: train[:20%]
    split: train
    split_thinking: true
    chat_template: qwen3
    field_messages: messages
    message_property_mappings:
      role: role
      content: content

      
val_set_size: 0.075
output_dir: ./outputs/out2
dataset_prepared_path: last_run_prepared

# sequence_len: 2048
sequence_len: 3072
sample_packing: true
eval_sample_packing: true


load_in_4bit: true
adapter: qlora
# lora_r: 16
# lora_alpha: 32

# lora_r: 32
# lora_alpha: 64


lora_r: 64
lora_alpha: 128

lora_target_modules:
  - q_proj
  - k_proj
  - v_proj
  - o_proj
  - down_proj
  - up_proj
lora_mlp_kernel: true
lora_qkv_kernel: true
lora_o_kernel: true

wandb_project:
wandb_entity:
wandb_watch:
wandb_name:
wandb_log_model:

gradient_accumulation_steps: 2
micro_batch_size: 4
num_epochs: 30
optimizer: adamw_torch_4bit
lr_scheduler: cosine
learning_rate: 0.00002

bf16: auto
tf32: true

gradient_checkpointing: offload
gradient_checkpointing_kwargs:
  use_reentrant: false
resume_from_checkpoint:
logging_steps: 1
flash_attention: true

warmup_ratio: 0.1
evals_per_epoch: 4
saves_per_epoch: 1
weight_decay: 0.01
special_tokens:

# save_first_step: true  # uncomment this to validate checkpoint saving works with your config

</details><br>

Qwen3-8B-Private

This model is a fine-tuned version of Qwen/Qwen3-8B on the trendmicro-ailab/Primus-Reasoning dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.9940
  • —Memory/max Active (gib): 8.07
  • —Memory/max Allocated (gib): 8.07
  • —Memory/device Reserved (gib): 10.78

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 2e-05
  • —trainbatchsize: 4
  • —evalbatchsize: 4
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 8
  • —optimizer: Use OptimizerNames.ADAMWTORCH4BIT with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 471
  • —training_steps: 4710

Training results

Training LossEpochStepValidation LossActive (gib)Allocated (gib)Reserved (gib)
No log001.27427.97.910.87
1.21020.2540401.23368.078.0710.76
1.01790.5079801.03968.078.0710.78
1.02780.76191200.96008.078.0710.78
0.97191.01271600.90878.078.0710.76
0.90521.26672000.86658.078.0710.76
0.82551.52062400.82748.078.0710.76
0.8131.77462800.79708.078.0710.76
0.8192.02543200.77578.078.0710.76
0.81382.27943600.75708.078.0710.78
0.77452.53334000.74338.078.0710.78
0.75872.78734400.73178.078.0710.78
0.72223.03814800.72248.078.0710.78
0.70923.29215200.71278.078.0710.76
0.66153.54605600.70718.078.0710.76
0.7153.86000.70268.078.0710.76
0.67474.05086400.69958.078.0710.78
0.70124.30486800.69398.078.0710.78
0.6984.55877200.69218.078.0710.78
0.65914.81277600.68748.078.0710.78
0.67165.06358000.68548.078.0710.78
0.70775.31758400.68588.078.0710.78
0.68175.57148800.68228.078.0710.78
0.66685.82549200.68008.078.0710.78
0.65436.07629600.68198.078.0710.78
0.63786.330210000.67988.078.0710.78
0.59226.584110400.67798.078.0710.78
0.6376.838110800.67738.078.0710.78
0.64787.088911200.67688.078.0710.78
0.64297.342911600.67818.078.0710.76
0.58477.596812000.67778.078.0710.76
0.64237.850812400.67348.078.0710.76
0.57938.101612800.67888.078.0710.76
0.57068.355613200.68028.078.0710.78
0.57298.609513600.67708.078.0710.78
0.67578.863514000.67558.078.0710.78
0.56439.114314400.68068.078.0710.78
0.53919.368314800.68258.078.0710.76
0.55659.622215200.68298.078.0710.76
0.59319.876215600.67778.078.0710.76
0.560810.127016000.68638.078.0710.78
0.563510.381016400.68648.078.0710.78
0.537910.634916800.68358.078.0710.78
0.543610.888917200.68588.078.0710.78
0.551111.139717600.69448.078.0710.78
0.53611.393718000.69528.078.0710.76
0.537111.647618400.69528.078.0710.76
0.576311.901618800.69108.078.0710.76
0.580212.152419200.70538.078.0710.78
0.580212.406319600.70628.078.0710.78
0.542212.660320000.70288.078.0710.78
0.47812.914320400.70278.078.0710.78
0.546713.165120800.72078.078.0710.78
0.534513.419021200.71828.078.0710.76
0.492213.673021600.71698.078.0710.76
0.506213.927022000.71658.078.0710.76
0.479714.177822400.73698.078.0710.76
0.443814.431722800.73358.078.0710.78
0.472614.685723200.72938.078.0710.78
0.465114.939723600.73058.078.0710.78
0.448915.190524000.75808.078.0710.78
0.444715.444424400.74948.078.0710.78
0.502715.698424800.74818.078.0710.78
0.488315.952425200.75048.078.0710.78
0.422316.203225600.76778.078.0710.78
0.49216.457126000.76888.078.0710.78
0.454116.711126400.77308.078.0710.78
0.480116.965126800.76888.078.0710.78
0.393217.215927200.79818.078.0710.78
0.420917.469827600.79008.078.0710.76
0.389117.723828000.79388.078.0710.76
0.415517.977828400.79038.078.0710.76
0.34718.228628800.82338.078.0710.78
0.355818.482529200.81728.078.0710.78
0.436518.736529600.82308.078.0710.78
0.445118.990530000.81818.078.0710.78
0.362719.241330400.85688.078.0710.78
0.33719.495230800.84038.078.0710.78
0.409419.749231200.84268.078.0710.78
0.422520.031600.83328.078.0710.78
0.348120.254032000.87528.078.0710.74
0.394720.507932400.87008.078.0710.78
0.410620.761932800.86498.078.0710.76
0.333321.012733200.87308.078.0710.78
0.355821.266733600.89668.078.0710.78
0.362521.520634000.89128.078.0710.78
0.342921.774634400.89188.078.0710.78
0.359722.025434800.91148.078.0710.78
0.344522.279435200.92178.078.0710.78
0.336622.533335600.91768.078.0710.78
0.355722.787336000.91818.078.0710.78
0.393723.038136400.93938.078.0710.78
0.316123.292136800.93918.078.0710.78
0.327223.546037200.94138.078.0710.78
0.375523.837600.93788.078.0710.78
0.296624.050838000.95648.078.0710.78
0.263924.304838400.95918.078.0710.78
0.30624.558738800.95998.078.0710.78
0.321524.812739200.95818.078.0710.78
0.39225.063539600.97138.078.0710.78
0.349425.317540000.97948.078.0710.78
0.324525.571440400.97208.078.0710.78
0.305325.825440800.97248.078.0710.78
0.30426.076241200.98498.078.0710.78
0.304326.330241600.98438.078.0710.76
0.342626.584142000.98468.078.0710.76
0.297926.838142400.98538.078.0710.76
0.352627.088942800.99128.078.0710.78
0.309527.342943200.98858.078.0710.78
0.298327.596843600.98988.078.0710.78
0.308627.850844000.99128.078.0710.78
0.319428.101644400.99198.078.0710.78
0.248428.355644800.99478.078.0710.78
0.345828.609545200.99438.078.0710.78
0.346728.863545600.99458.078.0710.78
0.329529.114346000.99368.078.0710.78
0.355529.368346400.99428.078.0710.78
0.327329.622246800.99408.078.0710.78

Framework versions

  • —PEFT 0.17.1
  • —Transformers 4.56.1
  • —Pytorch 2.7.1+cu126
  • —Datasets 4.0.0
  • —Tokenizers 0.22.1