CoolFace
Modelpublic

RichardErkhov/tsavage68_-_IE_L3_1000steps_1e6rate_05beta_cSFTDPO-gguf

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes258downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

IEL31000steps1e6rate05beta_cSFTDPO - GGUF

  • —Model creator: https://huggingface.co/tsavage68/
  • —Original model: https://huggingface.co/tsavage68/IEL31000steps1e6rate05beta_cSFTDPO/

Original model description: --- libraryname: transformers license: llama3 basemodel: tsavage68/IEL31000steps1e6rateSFT tags:

  • —trl
  • —dpo
  • —generatedfromtrainer model-index:
  • —name: IEL31000steps1e6rate05beta_cSFTDPO results: [] ---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

IEL31000steps1e6rate05beta_cSFTDPO

This model is a fine-tuned version of tsavage68/IE_L3_1000steps_1e6rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.1802
  • —Rewards/chosen: -1.4168
  • —Rewards/rejected: -13.8543
  • —Rewards/accuracies: 0.7400
  • —Rewards/margins: 12.4374
  • —Logps/rejected: -103.3358
  • —Logps/chosen: -85.6314
  • —Logits/rejected: -0.7970
  • —Logits/chosen: -0.7188

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-06
  • —trainbatchsize: 2
  • —evalbatchsize: 1
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 4
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 100
  • —training_steps: 1000

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.19060.4500.1802-1.0109-11.19030.740010.1794-98.0078-84.8196-0.7939-0.7206
0.13860.81000.1802-1.2190-12.16250.740010.9435-99.9523-85.2358-0.7944-0.7197
0.13861.21500.1802-1.2782-12.58520.740011.3070-100.7976-85.3541-0.7943-0.7189
0.17331.62000.1802-1.3094-13.02960.740011.7202-101.6864-85.4166-0.7948-0.7186
0.22532.02500.1802-1.3248-13.16250.740011.8377-101.9522-85.4473-0.7952-0.7186
0.13862.43000.1802-1.3337-13.26220.740011.9285-102.1515-85.4652-0.7942-0.7174
0.12132.83500.1802-1.3670-13.45070.740012.0837-102.5286-85.5317-0.7953-0.7178
0.19063.24000.1802-1.3818-13.53340.740012.1517-102.6941-85.5613-0.7964-0.7189
0.19063.64500.1802-1.3800-13.58990.740012.2099-102.8071-85.5577-0.7964-0.7189
0.20794.05000.1802-1.3816-13.67220.740012.2906-102.9716-85.5610-0.7966-0.7187
0.1564.45500.1802-1.4142-13.78000.740012.3657-103.1872-85.6262-0.7956-0.7175
0.12134.86000.1802-1.3864-13.77360.740012.3872-103.1744-85.5705-0.7974-0.7192
0.19065.26500.1802-1.4252-13.84500.740012.4197-103.3172-85.6483-0.7969-0.7187
0.24265.67000.1802-1.4087-13.81540.740012.4068-103.2581-85.6151-0.7974-0.7196
0.25996.07500.1802-1.4077-13.87120.740012.4635-103.3696-85.6131-0.7977-0.7194
0.12136.48000.1802-1.4158-13.90340.740012.4876-103.4339-85.6293-0.7977-0.7195
0.24266.88500.1802-1.4105-13.89220.740012.4817-103.4116-85.6187-0.7979-0.7200
0.17337.29000.1802-1.4075-13.86570.740012.4582-103.3587-85.6128-0.7970-0.7189
0.13867.69500.1802-1.4138-13.85230.740012.4386-103.3319-85.6253-0.7971-0.7188
0.1568.010000.1802-1.4168-13.85430.740012.4374-103.3358-85.6314-0.7970-0.7188

Framework versions

  • —Transformers 4.44.2
  • —Pytorch 2.0.0+cu117
  • —Datasets 3.0.0
  • —Tokenizers 0.19.1