CoolFace
Modelpublic

tsavage68/Transaminitis_L3_125steps_1e5rate_05beta_CSFTDPO

sourceHugging Facellama3updated 2y agoView on Hugging Face
0likes12downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

TransaminitisL3125steps1e5rate05beta_CSFTDPO

This model is a fine-tuned version of tsavage68/Transaminitis_L3_1000rate_1e7_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.7424
  • —Rewards/chosen: -16.3012
  • —Rewards/rejected: -16.2112
  • —Rewards/accuracies: 0.3500
  • —Rewards/margins: -0.0899
  • —Logps/rejected: -50.9772
  • —Logps/chosen: -51.1365
  • —Logits/rejected: -1.0737
  • —Logits/chosen: -1.0737

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-05
  • —trainbatchsize: 2
  • —evalbatchsize: 1
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 4
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 100
  • —training_steps: 125

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
1.33080.2251.4218-5.1457-5.29610.54000.1503-29.1468-28.8257-0.7892-0.7880
1.14980.4500.7304-4.8999-4.84250.4000-0.0574-28.2397-28.3340-2.1796-2.1797
1.28320.6750.9255-1.6896-4.28190.63002.5923-27.1184-21.9133-1.0885-1.0850
2.87640.81003.8444-19.0391-19.60420.54000.5651-57.7631-56.6124-0.1327-0.1327
0.77671.01250.7424-16.3012-16.21120.3500-0.0899-50.9772-51.1365-1.0737-1.0737

Framework versions

  • —Transformers 4.40.2
  • —Pytorch 2.0.0+cu117
  • —Datasets 2.19.1
  • —Tokenizers 0.19.1