CoolFace
Modelpublic

mhhmm/typescript-instruction-20k-v3

sourceHugging Facellama2updated 3y agoView on Hugging Face
0likes6downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

<img src="https://raw.githubusercontent.com/OpenAccess-AI-Collective/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl" width="200" height="32"/>

lora-out

This model is a fine-tuned version of codellama/CodeLlama-13b-Instruct-hf on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 0.4198

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0002
  • trainbatchsize: 8
  • evalbatchsize: 8
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 2
  • totaltrainbatch_size: 16
  • totalevalbatch_size: 16
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lrschedulertype: cosine
  • lrschedulerwarmup_steps: 10
  • num_epochs: 1

Training results

Training LossEpochStepValidation Loss
0.65550.0110.6550
0.65320.0570.6035
0.52780.1140.4977
0.54660.15210.4736
0.48320.2280.4637
0.50690.25350.4492
0.48640.3420.4436
0.46250.35490.4379
0.47920.4560.4336
0.46080.45630.4302
0.47380.5700.4266
0.48390.55770.4245
0.47910.6840.4227
0.47010.65910.4233
0.46120.7980.4225
0.44190.751050.4212
0.47050.81120.4199
0.44220.851190.4198
0.48890.91260.4198
0.49140.951330.4195
0.47991.01400.4198

Framework versions

  • Transformers 4.36.0.dev0
  • Pytorch 2.0.1+cu118
  • Datasets 2.15.0
  • Tokenizers 0.15.0

Training procedure

Framework versions

  • PEFT 0.6.0