CoolFace
Modelpublic

weijie210/zephyr-7b-dpo-reference

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes15downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

zephyr-7b-dpo-reference

This model is a fine-tuned version of alignment-handbook/zephyr-7b-sft-full on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 0.0429
  • Rewards/chosen: -0.5987
  • Rewards/rejected: -10.2552
  • Rewards/accuracies: 0.9741
  • Rewards/margins: 9.6565
  • Logps/rejected: -175.1052
  • Logps/chosen: -304.8906
  • Logits/rejected: -1.9643
  • Logits/chosen: -2.1592

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-07
  • trainbatchsize: 8
  • evalbatchsize: 8
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 4
  • totaltrainbatch_size: 32
  • totalevalbatch_size: 32
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lrschedulertype: linear
  • lrschedulerwarmup_ratio: 0.1
  • num_epochs: 1

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.08110.295000.0543-0.5167-8.91280.97208.3961-161.6813-304.0705-2.0037-2.1857
0.03620.5710000.0483-0.4980-9.58240.97209.0844-168.3771-303.8834-2.0113-2.2030
0.03180.8615000.0442-0.8458-10.59870.97209.7529-178.5403-307.3617-1.9506-2.1461

Framework versions

  • Transformers 4.36.1
  • Pytorch 2.0.1+cu117
  • Datasets 2.16.1
  • Tokenizers 0.15.0