CoolFace
Modelpublic

CharlesLi/OpenELM-1_1B-DPO-full-max-random-reward

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes19downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

OpenELM-1_1B-DPO-full-max-random-reward

This model was trained from scratch on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 194.2460
  • —Rewards/chosen: -660.0
  • —Rewards/rejected: -568.0
  • —Rewards/accuracies: 0.4277
  • —Rewards/margins: -89.5
  • —Logps/rejected: -57344.0
  • —Logps/chosen: -66560.0
  • —Logits/rejected: 7.5
  • —Logits/chosen: 7.0

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 5e-05
  • —trainbatchsize: 8
  • —evalbatchsize: 16
  • —seed: 42
  • —distributed_type: multi-GPU
  • —num_devices: 4
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 64
  • —totalevalbatch_size: 64
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_ratio: 0.1
  • —num_epochs: 3

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.69140.10471000.6983-0.3262-0.32230.4004-0.0052-320.0-352.0-9.4375-9.8125
0.69140.209420058.1418-201.0-173.00.4453-28.0-17664.0-20480.00.29690.2178
0.69140.314130097.2609-330.0-284.00.4258-44.75-28800.0-33280.0-0.3262-0.3027
0.69140.4188400102.8539-348.0-300.00.4297-47.75-30464.0-35328.0-0.7656-0.7461
0.69140.5236500108.8187-368.0-318.00.4277-50.25-32128.0-37120.0-0.2490-0.2812
0.69140.6283600114.7604-388.0-336.00.4355-53.0-33792.0-39168.01.251.1016
0.69140.7330700120.9475-410.0-354.00.4277-55.75-35584.0-41216.02.39062.1875
0.69140.8377800127.3012-432.0-372.00.4336-58.75-37632.0-43520.04.28123.9062
0.69140.9424900133.8314-454.0-392.00.4297-62.0-39424.0-45824.04.43754.0938
0.69141.04711000140.0195-476.0-410.00.4355-64.5-41472.0-47872.05.8755.4062
0.69141.15181100146.3645-496.0-430.00.4316-67.5-43264.0-49920.05.78125.375
0.69141.25651200151.9910-516.0-446.00.4336-70.0-44800.0-51968.06.3755.9375
0.69141.36131300157.8106-536.0-462.00.4297-73.0-46592.0-54016.07.06.5
0.69141.46601400163.0493-552.0-478.00.4316-75.5-48128.0-55552.07.34386.8125
0.69141.57071500168.1114-572.0-494.00.4277-77.5-49664.0-57344.07.28126.75
0.69141.67541600172.7765-588.0-506.00.4316-80.0-50944.0-58880.07.06256.5938
0.69141.78011700176.9677-600.0-520.00.4395-81.5-52224.0-60416.07.46886.9375
0.69141.88481800180.6313-612.0-532.00.4355-83.0-53248.0-61696.07.78127.25
0.69141.98951900183.7843-624.0-540.00.4258-84.5-54272.0-62720.07.6257.125
0.69142.09422000186.4619-632.0-548.00.4277-86.0-55040.0-63744.07.65627.125
0.69142.19902100188.7695-640.0-552.00.4258-87.0-55808.0-64512.07.59387.125
0.69142.30372200190.4722-648.0-560.00.4355-87.5-56320.0-65024.07.56257.0625
0.69142.40842300191.8555-652.0-564.00.4258-88.5-56576.0-65536.07.57.0312
0.69142.51312400192.9321-656.0-564.00.4258-89.0-56832.0-66048.07.43756.9688
0.69142.61782500193.6570-656.0-568.00.4258-89.0-57088.0-66048.07.46887.0
0.69142.72252600193.9604-660.0-568.00.4238-89.5-57344.0-66048.07.53127.0625
0.69142.82722700194.1360-660.0-568.00.4258-89.5-57344.0-66048.07.57.0312
0.69142.93192800194.2460-660.0-568.00.4277-89.5-57344.0-66560.07.57.0

Framework versions

  • —Transformers 4.44.2
  • —Pytorch 2.3.0
  • —Datasets 2.21.0
  • —Tokenizers 0.19.1