CoolFace
Modelpublic

CharlesLi/OpenELM-1_1B-DPO-full-random-pair

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes15downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

OpenELM-1_1B-DPO-full-random-pair

This model was trained from scratch on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 164.4180
  • Rewards/chosen: -560.0
  • Rewards/rejected: -482.0
  • Rewards/accuracies: 0.4277
  • Rewards/margins: -76.0
  • Logps/rejected: -48640.0
  • Logps/chosen: -56320.0
  • Logits/rejected: 3.1562
  • Logits/chosen: 2.7031

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • trainbatchsize: 8
  • evalbatchsize: 16
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 4
  • gradientaccumulationsteps: 2
  • totaltrainbatch_size: 64
  • totalevalbatch_size: 64
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lrschedulertype: cosine
  • lrschedulerwarmup_ratio: 0.1
  • num_epochs: 3

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.69140.10471000.6966-0.2227-0.21970.4570-0.0034-310.0-340.0-9.875-10.25
0.69140.20942004.2809-11.1875-10.3750.4316-0.8555-1328.0-1440.0-8.6875-9.1875
0.69140.314130095.7161-324.0-280.00.4258-44.0-28416.0-32768.0-2.5-2.5156
0.69140.418840099.5534-338.0-292.00.4277-46.0-29440.0-34048.0-2.625-2.6406
0.69140.5236500103.5082-352.0-304.00.4277-48.0-30592.0-35328.0-1.9688-2.0469
0.69140.6283600107.6879-366.0-316.00.4316-50.0-31872.0-36864.0-1.1328-1.2656
0.69140.7330700111.8930-380.0-328.00.4297-52.0-33024.0-38400.0-0.5117-0.7031
0.69140.8377800116.2988-394.0-340.00.4355-54.0-34304.0-39680.01.39061.0781
0.69140.9424900120.7803-410.0-354.00.4316-55.75-35584.0-41216.01.90621.5391
0.69141.04711000125.1435-424.0-366.00.4355-57.75-36864.0-42752.04.46883.9062
0.69141.15181100129.3826-440.0-380.00.4316-59.75-38144.0-44288.04.03123.5156
0.69141.25651200133.6557-454.0-392.00.4297-62.0-39424.0-45568.03.84383.2188
0.69141.36131300137.5098-466.0-404.00.4355-63.5-40704.0-46848.02.10941.7891
0.69141.46601400141.6271-482.0-416.00.4355-65.5-41728.0-48384.03.23442.7656
0.69141.57071500145.0692-492.0-426.00.4336-67.0-42752.0-49664.03.48442.9844
0.69141.67541600148.4839-504.0-436.00.4297-68.5-43776.0-50688.03.32812.8594
0.69141.78011700151.1965-512.0-444.00.4316-69.5-44544.0-51712.03.70313.2188
0.69141.88481800154.0215-524.0-452.00.4336-71.0-45568.0-52736.04.21883.6875
0.69141.98951900156.4897-532.0-460.00.4316-72.5-46080.0-53504.03.31252.875
0.69142.09422000158.3665-540.0-466.00.4336-73.0-46848.0-54016.03.18752.75
0.69142.19902100160.3225-544.0-470.00.4297-74.0-47360.0-54784.03.34382.8906
0.69142.30372200161.6044-548.0-474.00.4316-74.5-47616.0-55296.02.73442.3594
0.69142.40842300162.5378-552.0-478.00.4316-75.0-48128.0-55552.02.82812.4062
0.69142.51312400163.3184-556.0-480.00.4336-75.5-48128.0-55808.03.02.5469
0.69142.61782500163.9196-556.0-482.00.4316-75.5-48384.0-56064.03.18752.75
0.69142.72252600164.2697-556.0-482.00.4297-76.0-48640.0-56064.03.17192.7344
0.69142.82722700164.3540-560.0-482.00.4297-76.0-48640.0-56064.03.15622.7188
0.69142.93192800164.4180-560.0-482.00.4277-76.0-48640.0-56320.03.15622.7031

Framework versions

  • Transformers 4.44.2
  • Pytorch 2.3.0
  • Datasets 2.21.0
  • Tokenizers 0.19.1