ccjh/widthIrplus_msg_qm9s_mb_msd
017
<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->
widthIrplusmsgqm9smbmsd
This model is a fine-tuned version of /data/group/project1/models/Qwen3-32B on the QM9Sirtrain, the QM9Suvtrain, the QM9Sramantrain, the QM9Salltrain, the MBtrain, the MSGtrain, the MSDmstrain, the MSDnmrtrain, the MSDirtrain and the MSDalltrain datasets.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 5e-05
- trainbatchsize: 8
- evalbatchsize: 8
- seed: 42
- distributed_type: multi-GPU
- num_devices: 8
- gradientaccumulationsteps: 4
- totaltrainbatch_size: 256
- totalevalbatch_size: 64
- optimizer: Use adamwtorch with betas=(0.9,0.999) and epsilon=1e-08 and optimizerargs=No additional optimizer arguments
- lrschedulertype: cosine
- num_epochs: 1.0
Training results
Framework versions
- PEFT 0.15.1
- Transformers 4.51.3
- Pytorch 2.7.0+cu126
- Datasets 3.5.0
- Tokenizers 0.21.1
