CoolFace
Modelpublic

htlou/mm-interp-AA_preference_cocour_0_50

sourceHugging Faceotherupdated 2y agoView on Hugging Face
0likes8downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

AApreferencecocour050

This model is a fine-tuned version of llava-hf/llava-v1.6-mistral-7b-hf on the AApreferencecocour050 dataset. It achieves the following results on the evaluation set:

  • Loss: 0.4988
  • Rewards/chosen: 0.9568
  • Rewards/rejected: -1.9780
  • Rewards/accuracies: 0.8500
  • Rewards/margins: 2.9348
  • Logps/rejected: -217.8485
  • Logps/chosen: -263.8239
  • Logits/rejected: -2.1050
  • Logits/chosen: -2.1590

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-06
  • trainbatchsize: 8
  • evalbatchsize: 8
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 8
  • gradientaccumulationsteps: 4
  • totaltrainbatch_size: 256
  • totalevalbatch_size: 64
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lrschedulertype: cosine
  • lrschedulerwarmup_steps: 10
  • num_epochs: 3.0

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.60270.7463500.54911.3803-0.29830.82081.6786-201.0510-259.5887-2.4356-2.4575
0.27951.49251000.51121.1590-1.50930.84172.6683-213.1614-261.8016-1.9102-1.9756
0.15572.23881500.50331.3754-1.33250.85832.7079-211.3931-259.6372-2.1170-2.1696
0.13382.98512000.49830.9563-1.97620.85002.9325-217.8308-263.8291-2.1047-2.1588

Framework versions

  • Transformers 4.45.2
  • Pytorch 2.4.0+cu121
  • Datasets 2.21.0
  • Tokenizers 0.20.3