CoolFace
Modelpublic

tsavage68/UTI_L3_1000steps_1e5rate_03beta_CSFTDPO

sourceHugging Facellama3updated 2y agoView on Hugging Face
0likes15downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

UTIL31000steps1e5rate03beta_CSFTDPO

This model is a fine-tuned version of tsavage68/UTI_L3_1000steps_1e5rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.0069
  • —Rewards/chosen: 2.2757
  • —Rewards/rejected: -15.6836
  • —Rewards/accuracies: 0.9900
  • —Rewards/margins: 17.9593
  • —Logps/rejected: -115.4733
  • —Logps/chosen: -24.8934
  • —Logits/rejected: -1.4719
  • —Logits/chosen: -1.4307

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-05
  • —trainbatchsize: 2
  • —evalbatchsize: 1
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 4
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 100
  • —training_steps: 1000

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.00.6667500.00721.8402-13.35900.990015.1992-107.7247-26.3451-1.4305-1.3941
0.01731.33331000.00711.8455-14.40510.990016.2506-111.2116-26.3273-1.4331-1.3960
0.03472.01500.00692.3483-14.90500.990017.2533-112.8780-24.6513-1.4557-1.4154
0.02.66672000.00692.3179-15.01600.990017.3339-113.2480-24.7526-1.4584-1.4180
0.01733.33332500.00692.3120-15.08510.990017.3971-113.4783-24.7723-1.4616-1.4212
0.03474.03000.00692.3109-15.11440.990017.4254-113.5761-24.7759-1.4624-1.4219
0.01734.66673500.00692.3085-15.18590.990017.4944-113.8144-24.7841-1.4649-1.4242
0.01735.33334000.00692.2984-15.25710.990017.5555-114.0517-24.8176-1.4668-1.4260
0.01736.04500.00692.2945-15.34670.990017.6412-114.3504-24.8307-1.4680-1.4272
0.03476.66675000.00692.2859-15.42950.990017.7154-114.6264-24.8593-1.4694-1.4284
0.07.33335500.00692.2833-15.50570.990017.7890-114.8804-24.8681-1.4703-1.4293
0.03478.06000.00692.2775-15.57620.990017.8538-115.1155-24.8872-1.4709-1.4298
0.08.66676500.00692.2759-15.62060.990017.8965-115.2633-24.8928-1.4712-1.4301
0.01739.33337000.00692.2757-15.64250.990017.9182-115.3363-24.8933-1.4714-1.4302
0.010.07500.00692.2743-15.66500.990017.9392-115.4112-24.8982-1.4717-1.4305
0.017310.66678000.00692.2739-15.67850.990017.9524-115.4563-24.8992-1.4719-1.4307
0.011.33338500.00692.2703-15.66670.990017.9370-115.4169-24.9113-1.4717-1.4306
0.012.09000.00692.2749-15.67710.990017.9520-115.4516-24.8959-1.4719-1.4307
0.017312.66679500.00692.2732-15.67530.990017.9485-115.4458-24.9018-1.4719-1.4307
0.013.333310000.00692.2757-15.68360.990017.9593-115.4733-24.8934-1.4719-1.4307

Framework versions

  • —Transformers 4.41.2
  • —Pytorch 2.0.0+cu117
  • —Datasets 2.19.2
  • —Tokenizers 0.19.1