CoolFace
Modelpublic

RichardErkhov/tsavage68_-_Na_M2_1000steps_1e7rate_05beta_cSFTDPO-gguf

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes162downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

NaM21000steps1e7rate05beta_cSFTDPO - GGUF

  • —Model creator: https://huggingface.co/tsavage68/
  • —Original model: https://huggingface.co/tsavage68/NaM21000steps1e7rate05beta_cSFTDPO/

Original model description: --- libraryname: transformers license: apache-2.0 basemodel: tsavage68/NaM21000steps1e7SFT tags:

  • —trl
  • —dpo
  • —generatedfromtrainer model-index:
  • —name: NaM21000steps1e7rate05beta_cSFTDPO results: [] ---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

NaM21000steps1e7rate05beta_cSFTDPO

This model is a fine-tuned version of tsavage68/Na_M2_1000steps_1e7_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.0000
  • —Rewards/chosen: 3.4353
  • —Rewards/rejected: -12.0460
  • —Rewards/accuracies: 1.0
  • —Rewards/margins: 15.4813
  • —Logps/rejected: -104.0153
  • —Logps/chosen: -41.2618
  • —Logits/rejected: -2.5171
  • —Logits/chosen: -2.5312

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-07
  • —trainbatchsize: 2
  • —evalbatchsize: 1
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 4
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 100
  • —training_steps: 1000

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.00.2667500.00002.4333-8.79461.011.2279-97.5125-43.2658-2.5259-2.5391
0.00.53331000.00002.7977-9.99361.012.7913-99.9105-42.5369-2.5223-2.5359
0.00.81500.00002.9419-10.65511.013.5970-101.2335-42.2486-2.5210-2.5347
0.01.06672000.00003.0397-10.99891.014.0386-101.9212-42.0530-2.5209-2.5347
0.01.33332500.00003.1479-11.23651.014.3844-102.3963-41.8365-2.5209-2.5348
0.01.63000.00003.1788-11.46041.014.6393-102.8442-41.7747-2.5197-2.5337
0.01.86673500.00003.2803-11.63061.014.9109-103.1846-41.5718-2.5199-2.5339
0.02.13334000.00003.3009-11.78681.015.0878-103.4970-41.5305-2.5189-2.5328
0.02.44500.00003.3596-11.86641.015.2260-103.6562-41.4132-2.5179-2.5319
0.02.66675000.00003.3481-11.93381.015.2818-103.7909-41.4363-2.5176-2.5316
0.02.93335500.00003.3954-11.95911.015.3545-103.8415-41.3415-2.5186-2.5326
0.03.26000.00003.4233-12.04361.015.4669-104.0106-41.2858-2.5181-2.5321
0.03.46676500.00003.4170-12.05351.015.4704-104.0303-41.2985-2.5183-2.5323
0.03.73337000.00003.3924-12.07361.015.4660-104.0705-41.3476-2.5178-2.5318
0.04.07500.00003.4428-12.05661.015.4994-104.0365-41.2468-2.5180-2.5321
0.04.26678000.00003.4331-12.04691.015.4800-104.0172-41.2661-2.5173-2.5314
0.04.53338500.00003.4177-12.07941.015.4970-104.0821-41.2971-2.5172-2.5312
0.04.89000.00003.4353-12.04601.015.4813-104.0153-41.2618-2.5171-2.5312
0.05.06679500.00003.4353-12.04601.015.4813-104.0153-41.2618-2.5171-2.5312
0.05.333310000.00003.4353-12.04601.015.4813-104.0153-41.2618-2.5171-2.5312

Framework versions

  • —Transformers 4.44.2
  • —Pytorch 2.4.0+cu121
  • —Datasets 2.21.0
  • —Tokenizers 0.19.1