CoolFace
Modelpublic

RichardErkhov/CharlesLi_-_OpenELM-1_1B-DPO-full-max-reward-most-similar-gguf

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes597downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

OpenELM-1_1B-DPO-full-max-reward-most-similar - GGUF

  • —Model creator: https://huggingface.co/CharlesLi/
  • —Original model: https://huggingface.co/CharlesLi/OpenELM-1_1B-DPO-full-max-reward-most-similar/

Original model description: --- library_name: transformers tags:

  • —trl
  • —dpo
  • —alignment-handbook
  • —generatedfromtrainer model-index:
  • —name: OpenELM-1_1B-DPO-full-max-reward-most-similar results: [] ---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

OpenELM-1_1B-DPO-full-max-reward-most-similar

This model was trained from scratch on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 1.6465
  • —Rewards/chosen: -17.75
  • —Rewards/rejected: -19.75
  • —Rewards/accuracies: 0.6055
  • —Rewards/margins: 2.0469
  • —Logps/rejected: -2272.0
  • —Logps/chosen: -2096.0
  • —Logits/rejected: 2.0312
  • —Logits/chosen: 0.2393

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 5e-05
  • —trainbatchsize: 8
  • —evalbatchsize: 16
  • —seed: 42
  • —distributed_type: multi-GPU
  • —num_devices: 4
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 64
  • —totalevalbatch_size: 64
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_ratio: 0.1
  • —num_epochs: 3

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.57860.10471000.6689-1.8203-2.06250.60940.2373-494.0-500.0-9.25-9.75
0.53590.20942000.7366-3.3281-3.81250.58980.4824-672.0-652.0-1.7812-2.8906
0.51630.31413000.6974-4.25-4.84380.64260.6016-776.0-744.0-5.4688-6.75
0.51270.41884000.7937-5.375-6.06250.60160.6797-896.0-856.0-7.9375-9.125
0.50470.52365000.7909-4.5938-5.21880.57030.6523-812.0-776.0-3.7188-5.5938
0.50570.62836000.8288-5.375-6.1250.59180.7539-904.0-856.0-4.5-6.4062
0.480.73307000.7987-5.5312-6.40620.62890.8633-928.0-872.0-3.8438-5.6562
0.47510.83778000.8430-7.0625-7.78120.55860.7070-1064.0-1024.0-4.3125-6.125
0.44080.94249000.8971-8.3125-9.18750.59960.9023-1208.0-1152.0-6.3438-8.1875
0.16091.047110000.9796-8.1875-9.18750.59961.0156-1208.0-1136.0-1.7734-3.7656
0.15511.151811001.2334-13.8125-15.06250.59381.2422-1792.0-1704.0-0.2617-2.0312
0.15841.256512001.0642-10.375-11.56250.59181.1641-1440.0-1360.0-2.1875-3.9844
0.16181.361313000.9750-9.1875-10.31250.62111.1484-1320.0-1240.0-1.25-3.0781
0.16671.466014001.0401-9.75-11.1250.61911.3125-1400.0-1296.0-1.1094-3.1875
0.17141.570715001.0380-10.6875-12.06250.62301.3438-1496.0-1392.0-0.2578-2.1719
0.14061.675416001.0427-11.25-12.6250.62111.375-1552.0-1440.0-0.0874-2.0469
0.11951.780117001.1374-12.25-13.6250.61331.3906-1648.0-1544.0-0.4316-2.1875
0.12911.884818001.0742-11.6875-13.06250.59381.3438-1592.0-1488.00.0305-1.7344
0.12361.989519001.1539-13.0-14.3750.58401.3984-1728.0-1616.00.7383-0.9727
0.02642.094220001.5533-16.5-18.250.58401.75-2112.0-1968.01.1562-0.625
0.02222.199021001.6053-17.375-19.250.59571.8906-2224.0-2064.02.07810.3105
0.02662.303722001.5843-17.125-19.00.60551.8672-2192.0-2032.01.92970.0918
0.02472.408423001.6309-17.875-19.8750.60942.0-2288.0-2112.02.17190.3652
0.03812.513124001.6237-17.75-19.6250.60551.9219-2256.0-2096.02.00.2354
0.03072.617825001.6102-17.375-19.3750.60552.0156-2224.0-2064.01.91410.1069
0.02592.722526001.6399-17.75-19.750.60352.0469-2272.0-2096.02.04690.2773
0.02792.827227001.6252-17.5-19.50.60742.0312-2240.0-2064.01.96090.1533
0.02192.931928001.6465-17.75-19.750.60552.0469-2272.0-2096.02.03120.2393

Framework versions

  • —Transformers 4.45.1
  • —Pytorch 2.3.0
  • —Datasets 3.0.1
  • —Tokenizers 0.20.0