CoolFace
Modelpublic

RichardErkhov/tsavage68_-_IE_L3_1000steps_1e6rate_03beta_cSFTDPO-gguf

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes216downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

IEL31000steps1e6rate03beta_cSFTDPO - GGUF

  • —Model creator: https://huggingface.co/tsavage68/
  • —Original model: https://huggingface.co/tsavage68/IEL31000steps1e6rate03beta_cSFTDPO/

Original model description: --- libraryname: transformers license: llama3 basemodel: tsavage68/IEL31000steps1e6rateSFT tags:

  • —trl
  • —dpo
  • —generatedfromtrainer model-index:
  • —name: IEL31000steps1e6rate03beta_cSFTDPO results: [] ---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

IEL31000steps1e6rate03beta_cSFTDPO

This model is a fine-tuned version of tsavage68/IE_L3_1000steps_1e6rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.1802
  • —Rewards/chosen: -1.3199
  • —Rewards/rejected: -13.3530
  • —Rewards/accuracies: 0.7400
  • —Rewards/margins: 12.0331
  • —Logps/rejected: -120.1372
  • —Logps/chosen: -87.1973
  • —Logits/rejected: -0.8052
  • —Logits/chosen: -0.7124

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-06
  • —trainbatchsize: 2
  • —evalbatchsize: 1
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 4
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 100
  • —training_steps: 1000

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.19070.4500.1802-1.0923-10.46800.74009.3757-110.5205-86.4386-0.7963-0.7114
0.13860.81000.1802-1.2190-11.57160.740010.3526-114.1993-86.8611-0.7960-0.7088
0.13861.21500.1802-1.2269-11.87970.740010.6528-115.2263-86.8875-0.7973-0.7092
0.17331.62000.1802-1.2628-12.45620.740011.1934-117.1479-87.0072-0.7983-0.7088
0.22532.02500.1802-1.2811-12.61090.740011.3298-117.6637-87.0682-0.8005-0.7100
0.13862.43000.1802-1.2819-12.68210.740011.4002-117.9011-87.0709-0.8009-0.7104
0.12132.83500.1802-1.2857-12.92520.740011.6395-118.7114-87.0834-0.8024-0.7110
0.19063.24000.1802-1.2904-12.99290.740011.7024-118.9368-87.0992-0.8026-0.7109
0.19063.64500.1802-1.2935-13.03200.740011.7385-119.0673-87.1095-0.8030-0.7112
0.20794.05000.1802-1.3034-13.17280.740011.8694-119.5364-87.1423-0.8047-0.7126
0.1564.45500.1802-1.3085-13.22420.740011.9157-119.7078-87.1593-0.8035-0.7118
0.12134.86000.1802-1.2992-13.24110.740011.9418-119.7642-87.1285-0.8054-0.7131
0.19065.26500.1802-1.3144-13.31560.740012.0011-120.0125-87.1792-0.8048-0.7117
0.24265.67000.1802-1.2925-13.30310.740012.0106-119.9710-87.1061-0.8043-0.7117
0.25996.07500.1802-1.3084-13.32980.740012.0213-120.0597-87.1592-0.8052-0.7126
0.12136.48000.1802-1.3118-13.34770.740012.0359-120.1197-87.1704-0.8039-0.7116
0.24266.88500.1802-1.3228-13.36200.740012.0392-120.1673-87.2071-0.8052-0.7125
0.17337.29000.1802-1.3137-13.33790.740012.0242-120.0870-87.1768-0.8052-0.7125
0.13867.69500.1802-1.3070-13.35300.740012.0460-120.1374-87.1545-0.8053-0.7127
0.1568.010000.1802-1.3199-13.35300.740012.0331-120.1372-87.1973-0.8052-0.7124

Framework versions

  • —Transformers 4.44.2
  • —Pytorch 2.0.0+cu117
  • —Datasets 3.0.0
  • —Tokenizers 0.19.1