CoolFace
Modelpublic

asadfgglie/mDeBERTa-v3-base-xnli-multilingual-zeroshot-v1.1-seed20241201

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes36downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

mDeBERTa-v3-base-xnli-multilingual-zeroshot-v1.1-seed20241201

This model use same hyper-parameter with asadfgglie/mDeBERTa-v3-base-xnli-multilingual-zeroshot-v1.1, except RANDOM_SEED.

Original version use RANDOM_SEED=42, this version use RANDOM_SEED=20241201.

This model is a fine-tuned version of asadfgglie/mDeBERTa-v3-base-xnli-multilingual-zeroshot-v1.0 on the None dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.6134
  • —F1 Macro: 0.8616
  • —F1 Micro: 0.8634
  • —Accuracy Balanced: 0.8616
  • —Accuracy: 0.8634
  • —Precision Macro: 0.8616
  • —Recall Macro: 0.8616
  • —Precision Micro: 0.8634
  • —Recall Micro: 0.8634

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 2e-05
  • —trainbatchsize: 16
  • —evalbatchsize: 128
  • —seed: 20241201
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 32
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —lrschedulerwarmup_ratio: 0.06
  • —num_epochs: 3

Training results

Training LossEpochStepValidation LossF1 MacroF1 MicroAccuracy BalancedAccuracyPrecision MacroRecall MacroPrecision MicroRecall Micro
0.20340.172000.42410.84810.85180.84510.85180.85410.84510.85180.8518
0.2190.344000.41780.86080.86240.86150.86240.86020.86150.86240.8624
0.21420.516000.38100.85720.86020.85480.86020.86130.85480.86020.8602
0.1990.688000.43140.85370.85710.85080.85710.85900.85080.85710.8571
0.20050.8510000.42820.85720.86020.85470.86020.86150.85470.86020.8602
0.18461.0212000.46310.86910.87030.87070.87030.86810.87070.87030.8703
0.1541.1914000.49220.85990.86130.86100.86130.85900.86100.86130.8613
0.14321.3516000.50200.85400.85600.85400.85600.85410.85400.85600.8560
0.13351.5218000.53130.84790.85070.84610.85070.85050.84610.85070.8507
0.13731.6920000.50180.85460.85710.85330.85710.85630.85330.85710.8571
0.1281.8622000.48960.86440.86550.86650.86550.86330.86650.86550.8655
0.12572.0324000.49220.86480.86660.86480.86660.86480.86480.86660.8666
0.09592.226000.58140.85890.86130.85760.86130.86060.85760.86130.8613
0.09182.3728000.59870.86170.86340.86180.86340.86150.86180.86340.8634
0.09922.5430000.61170.86310.86500.86290.86500.86340.86290.86500.8650
0.08972.7132000.61910.85830.86020.85830.86020.85840.85830.86020.8602
0.10652.8834000.62210.86250.86450.86190.86450.86310.86190.86450.8645

Eval result

Datasetsasadfgglie/nli-zh-tw-all/testasadfgglie/BanBan_2024-10-17-facial_expressions-nli/testeval_datasettest_dataset
eval_loss0.5750.3560.6130.538
evalf1macro0.870.8960.8620.875
evalf1micro0.8710.8960.8630.876
evalaccuracybalanced0.8690.8960.8620.875
eval_accuracy0.8710.8960.8630.876
evalprecisionmacro0.8710.8980.8620.877
evalrecallmacro0.8690.8960.8620.875
evalprecisionmicro0.8710.8960.8630.876
evalrecallmicro0.8710.8960.8630.876
eval_runtime229.7324.24851.374204.44
evalsamplesper_second37.0222.6836.76936.964
evalstepsper_second0.2921.8830.2920.293
Size of dataset850094618897557

Framework versions

  • —Transformers 4.33.3
  • —Pytorch 2.5.1+cu121
  • —Datasets 2.14.7
  • —Tokenizers 0.13.3