CoolFace
Modelpublic

mariana-coelho-9/lora-adapter-Llama-3.1-8B-Instruct-bnb-4bit-icd-10

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes
Model Card

Uploaded model

  • —Developed by: mariana-coelho-9
  • —License: apache-2.0
  • —Finetuned from model : unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit

This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>

Dataset Information

Dataset used: mariana-coelho-9/icd-10

  • —Training set size: 28261
  • —Validation set size: 3141

Memory and Time Statistics

GPU = Tesla T4.

  • —Max memory = 14.748 GB.
  • —Peak reserved memory = 7.717 GB.
  • —Peak reserved memory for training = 0.0 GB.
  • —Peak reserved memory % of max memory = 52.326 %.
  • —Peak reserved memory for training % of max memory = 0.0 %.
  • —2033.5873 seconds used for training.
  • —33.89 minutes used for training.

Hyperparameters

Hyperparameters:

  • —'r': 16,
  • —'targetmodules': ['qproj', 'kproj', 'vproj', 'oproj', 'gateproj', 'upproj', 'downproj']
  • —'lora_alpha': 16
  • —'lora_dropout': 0
  • —'bias': 'none'
  • —'usegradientcheckpointing': 'unsloth'
  • —'random_state': 3407
  • —'use_rslora': False
  • —'loftq_config': None
  • —'perdevicetrainbatchsize': 2
  • —'gradientaccumulationsteps': 4
  • —'warmupsteps': 5, 'maxsteps': 60
  • —'learning_rate': 0.0002
  • —'fp16': True
  • —'bf16': False
  • —'logging_steps': 10
  • —'optim': 'adamw_8bit'
  • —'weight_decay': 0.01
  • —'lrschedulertype': 'linear'
  • —'seed': 3407