CoolFace
Modelpublic

rayonlabs/L-MChat-7b-Bitext-retail-ecommerce-llm-chatbot-training-dataset-95ac04d6-a803-4337-a004-46c74f84

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes7downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

<img src="https://raw.githubusercontent.com/axolotl-ai-cloud/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl" width="200" height="32"/> <br>

83f42180-15c1-490a-b36a-7a9624357888

This model is a fine-tuned version of Artples/L-MChat-7b on the None dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.4879

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 0.000208
  • —trainbatchsize: 4
  • —evalbatchsize: 4
  • —seed: 80
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 8
  • —optimizer: Use OptimizerNames.ADAMWBNB with betas=(0.9,0.999) and epsilon=1e-08 and optimizerargs=No additional optimizer arguments
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 50
  • —training_steps: 500

Training results

Training LossEpochStepValidation Loss
No log0.000211.0814
1.31890.0094500.8371
1.38530.01881000.7781
1.31050.02821500.7884
1.27170.03762000.7464
1.24350.04702500.7073
1.20820.05643000.6344
1.17660.06583500.5577
1.10590.07524000.5161
1.06280.08464500.4936
1.04530.09405000.4879

Framework versions

  • —PEFT 0.13.2
  • —Transformers 4.46.0
  • —Pytorch 2.5.0+cu124
  • —Datasets 3.0.1
  • —Tokenizers 0.20.1