CoolFace
Modelpublic

SotirisLegkas/Llama3_ALL_BCE_translations_19_shuffled_special_tokens

sourceHugging Facellama3updated 2y agoView on Hugging Face
0likes5downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

Llama3ALLBCEtranslations19shuffledspecial_tokens

This model is a fine-tuned version of meta-llama/Meta-Llama-3-8B-Instruct on the None dataset. It achieves the following results on the evaluation set:

  • —Loss: 1.4776
  • —F1 Macro 0.1: 0.0818
  • —F1 Macro 0.15: 0.0922
  • —F1 Macro 0.2: 0.1027
  • —F1 Macro 0.25: 0.1130
  • —F1 Macro 0.3: 0.1230
  • —F1 Macro 0.35: 0.1336
  • —F1 Macro 0.4: 0.1440
  • —F1 Macro 0.45: 0.1551
  • —F1 Macro 0.5: 0.1663
  • —F1 Macro 0.55: 0.1778
  • —F1 Macro 0.6: 0.1879
  • —F1 Macro 0.65: 0.1987
  • —F1 Macro 0.7: 0.2090
  • —F1 Macro 0.75: 0.2178
  • —F1 Macro 0.8: 0.2211
  • —F1 Macro 0.85: 0.2205
  • —F1 Macro 0.9: 0.2010
  • —F1 Macro 0.95: 0.1457
  • —Threshold 0: 0.65
  • —Threshold 1: 0.75
  • —Threshold 2: 0.7
  • —Threshold 3: 0.85
  • —Threshold 4: 0.8
  • —Threshold 5: 0.85
  • —Threshold 6: 0.8
  • —Threshold 7: 0.8
  • —Threshold 8: 0.85
  • —Threshold 9: 0.75
  • —Threshold 10: 0.85
  • —Threshold 11: 0.8
  • —Threshold 12: 0.85
  • —Threshold 13: 0.95
  • —Threshold 14: 0.85
  • —Threshold 15: 0.75
  • —Threshold 16: 0.85
  • —Threshold 17: 0.8
  • —Threshold 18: 0.9
  • —0: 0.0619
  • —1: 0.1388
  • —2: 0.1978
  • —3: 0.1328
  • —4: 0.2961
  • —5: 0.3489
  • —6: 0.3179
  • —7: 0.1268
  • —8: 0.2043
  • —9: 0.3668
  • —10: 0.3216
  • —11: 0.3669
  • —12: 0.1276
  • —13: 0.1205
  • —14: 0.2264
  • —15: 0.1576
  • —16: 0.3078
  • —17: 0.3722
  • —18: 0.125
  • —Max F1: 0.2211
  • —Mean F1: 0.2273

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 5e-06
  • —trainbatchsize: 8
  • —evalbatchsize: 8
  • —seed: 2024
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —lrschedulerwarmup_ratio: 0.1
  • —num_epochs: 4

Training results

Training LossEpochStepValidation LossF1 Macro 0.1F1 Macro 0.15F1 Macro 0.2F1 Macro 0.25F1 Macro 0.3F1 Macro 0.35F1 Macro 0.4F1 Macro 0.45F1 Macro 0.5F1 Macro 0.55F1 Macro 0.6F1 Macro 0.65F1 Macro 0.7F1 Macro 0.75F1 Macro 0.8F1 Macro 0.85F1 Macro 0.9F1 Macro 0.95Threshold 0Threshold 1Threshold 2Threshold 3Threshold 4Threshold 5Threshold 6Threshold 7Threshold 8Threshold 9Threshold 10Threshold 11Threshold 12Threshold 13Threshold 14Threshold 15Threshold 16Threshold 17Threshold 180123456789101112131415161718Max F1Mean F1
3.38241.055954.38470.07000.07610.08180.08770.09360.10000.10640.11340.11960.12650.13270.13810.14320.14830.14650.14170.12910.08360.650.90.850.90.750.60.80.750.90.90.90.850.90.00.850.750.60.60.90.06490.08790.16030.08990.25890.28760.26830.10360.12450.28560.23870.30330.07260.00.17790.11090.21920.27430.06410.14830.1680
2.48592.0111901.75370.08810.09940.11110.12100.13100.14010.14720.15410.16070.16760.16970.17310.17680.17610.17130.15750.13650.09270.550.70.850.80.40.350.950.750.70.850.80.650.80.950.80.70.850.60.750.05340.12410.19240.10200.27380.31630.30720.11090.17930.34140.28890.33320.08310.08700.21370.13050.28810.33960.12540.17680.2048
1.75613.0167851.46330.08400.09540.10620.11640.12710.13820.14850.15970.17130.18090.18950.19760.20560.21130.21150.19950.18050.11840.60.750.750.950.80.70.90.80.80.70.80.80.90.950.750.80.70.70.80.05810.13950.19460.12350.28180.33910.31510.12020.19970.36560.30560.36300.13400.10870.22720.14820.29530.35890.12330.21150.2211
1.27094.0223801.47760.08180.09220.10270.11300.12300.13360.14400.15510.16630.17780.18790.19870.20900.21780.22110.22050.20100.14570.650.750.70.850.80.850.80.80.850.750.850.80.850.950.850.750.850.80.90.06190.13880.19780.13280.29610.34890.31790.12680.20430.36680.32160.36690.12760.12050.22640.15760.30780.37220.1250.22110.2273

Framework versions

  • —PEFT 0.10.0
  • —Transformers 4.40.2
  • —Pytorch 2.2.2+cu121
  • —Datasets 2.18.0
  • —Tokenizers 0.19.1