CoolFace
Modelpublic

RichardErkhov/tsavage68_-_IE_L3_350steps_1e8rate_03beta_cSFTDPO-gguf

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes196downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

IEL3350steps1e8rate03beta_cSFTDPO - GGUF

  • —Model creator: https://huggingface.co/tsavage68/
  • —Original model: https://huggingface.co/tsavage68/IEL3350steps1e8rate03beta_cSFTDPO/

Original model description: --- libraryname: transformers license: llama3 basemodel: tsavage68/IEL31000steps1e6rateSFT tags:

  • —trl
  • —dpo
  • —generatedfromtrainer model-index:
  • —name: IEL3350steps1e8rate03beta_cSFTDPO results: [] ---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

IEL3350steps1e8rate03beta_cSFTDPO

This model is a fine-tuned version of tsavage68/IE_L3_1000steps_1e6rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.6896
  • —Rewards/chosen: -0.0071
  • —Rewards/rejected: -0.0198
  • —Rewards/accuracies: 0.4400
  • —Rewards/margins: 0.0127
  • —Logps/rejected: -75.6932
  • —Logps/chosen: -82.8214
  • —Logits/rejected: -0.7977
  • —Logits/chosen: -0.7408

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-08
  • —trainbatchsize: 2
  • —evalbatchsize: 1
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 4
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 100
  • —training_steps: 350

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.69120.4500.6940-0.0075-0.01040.40000.0029-75.6618-82.8226-0.7964-0.7393
0.69470.81000.69250.0014-0.00570.38500.0070-75.6461-82.7931-0.7963-0.7394
0.68811.21500.7003-0.0102-0.00200.375-0.0082-75.6340-82.8318-0.7969-0.7398
0.67761.62000.6938-0.0057-0.00980.3750.0041-75.6601-82.8168-0.7970-0.7399
0.68592.02500.6850-0.0033-0.02500.43500.0217-75.7105-82.8087-0.7975-0.7405
0.70242.43000.6893-0.0075-0.02070.44000.0132-75.6964-82.8228-0.7977-0.7408
0.68022.83500.6896-0.0071-0.01980.44000.0127-75.6932-82.8214-0.7977-0.7408

Framework versions

  • —Transformers 4.44.2
  • —Pytorch 2.0.0+cu117
  • —Datasets 3.0.0
  • —Tokenizers 0.19.1