CoolFace
Modelpublic

RichardErkhov/tsavage68_-_IE_M2_1000steps_1e6rate_01beta_cSFTDPO-gguf

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes246downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

IEM21000steps1e6rate01beta_cSFTDPO - GGUF

  • —Model creator: https://huggingface.co/tsavage68/
  • —Original model: https://huggingface.co/tsavage68/IEM21000steps1e6rate01beta_cSFTDPO/

Original model description: --- libraryname: transformers license: apache-2.0 basemodel: tsavage68/IEM21000steps1e7rateSFT tags:

  • —trl
  • —dpo
  • —generatedfromtrainer model-index:
  • —name: IEM21000steps1e6rate01beta_cSFTDPO results: [] ---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

IEM21000steps1e6rate01beta_cSFTDPO

This model is a fine-tuned version of tsavage68/IE_M2_1000steps_1e7rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.3743
  • —Rewards/chosen: -0.0096
  • —Rewards/rejected: -8.8855
  • —Rewards/accuracies: 0.4600
  • —Rewards/margins: 8.8759
  • —Logps/rejected: -129.8764
  • —Logps/chosen: -42.3012
  • —Logits/rejected: -2.8667
  • —Logits/chosen: -2.7910

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-06
  • —trainbatchsize: 2
  • —evalbatchsize: 1
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 4
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 100
  • —training_steps: 1000

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.45070.4500.3744-0.2335-4.59850.46004.3650-87.0066-44.5404-2.8777-2.8157
0.38120.81000.3743-0.5108-6.69210.46006.1813-107.9430-47.3138-2.8657-2.7955
0.31191.21500.3743-0.1626-7.31450.46007.1519-114.1667-43.8313-2.8595-2.7859
0.36391.62000.3743-0.0733-7.77210.46007.6988-118.7424-42.9385-2.8656-2.7905
0.43322.02500.3743-0.0463-8.04790.46008.0016-121.5008-42.6684-2.8656-2.7903
0.39862.43000.3743-0.0312-8.22410.46008.1929-123.2630-42.5179-2.8658-2.7905
0.39862.83500.3743-0.0173-8.33430.46008.3171-124.3653-42.3781-2.8664-2.7908
0.45053.24000.3743-0.0158-8.51770.46008.5019-126.1987-42.3632-2.8666-2.7910
0.45053.64500.3743-0.0135-8.55180.46008.5383-126.5393-42.3402-2.8666-2.7910
0.43324.05000.3743-0.0117-8.66420.46008.6525-127.6642-42.3228-2.8665-2.7909
0.32924.45500.3743-0.0128-8.69570.46008.6829-127.9786-42.3337-2.8666-2.7910
0.36394.86000.3743-0.0122-8.79910.46008.7869-129.0126-42.3276-2.8671-2.7915
0.45055.26500.3743-0.0110-8.83120.46008.8202-129.3338-42.3151-2.8667-2.7910
0.45055.67000.3743-0.0140-8.85230.46008.8383-129.5449-42.3457-2.8668-2.7911
0.36396.07500.3743-0.0142-8.87600.46008.8618-129.7817-42.3476-2.8666-2.7909
0.24266.48000.3743-0.0114-8.88480.46008.8734-129.8699-42.3197-2.8667-2.7910
0.50256.88500.3743-0.0110-8.88240.46008.8714-129.8454-42.3153-2.8666-2.7910
0.31197.29000.3743-0.0122-8.89320.46008.8810-129.9536-42.3276-2.8668-2.7911
0.34667.69500.3743-0.0106-8.88840.46008.8778-129.9054-42.3112-2.8667-2.7910
0.38128.010000.3743-0.0096-8.88550.46008.8759-129.8764-42.3012-2.8667-2.7910

Framework versions

  • —Transformers 4.44.2
  • —Pytorch 2.0.0+cu117
  • —Datasets 3.0.0
  • —Tokenizers 0.19.1