CoolFace
Modelpublic

RichardErkhov/DevQuasar_-_analytical_reasoning_r16a32_unsloth-Llama-3.2-3B-Instruct-bnb-4bit-4bits

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes11downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

analyticalreasoningr16a32_unsloth-Llama-3.2-3B-Instruct-bnb-4bit - bnb 4bits

  • —Model creator: https://huggingface.co/DevQuasar/
  • —Original model: https://huggingface.co/DevQuasar/analyticalreasoningr16a32_unsloth-Llama-3.2-3B-Instruct-bnb-4bit/

Original model description: --- base_model: unsloth/Llama-3.2-3B-Instruct-bnb-4bit datasets:

  • —microsoft/orca-agentinstruct-1M-v1 pipelinetag: text-generation libraryname: transformers license: llama3.2 tags:
  • —unsloth
  • —transformers model-index:
  • —name: analyticalreasoningr16a32_unsloth-Llama-3.2-3B-Instruct-bnb-4bit results:
  • —task: type: text-generation dataset: type: lm-evaluation-harness name: bbh metrics:
  • —name: accnorm type: accnorm value: 0.4168 verified: false
  • —task: type: text-generation dataset: type: lm-evaluation-harness name: gpqa metrics:
  • —name: accnorm type: accnorm value: 0.2691 verified: false
  • —task: type: text-generation dataset: type: lm-evaluation-harness name: math metrics:
  • —name: exactmatch type: exactmatch value: 0.0867 verified: false
  • —task: type: text-generation dataset: type: lm-evaluation-harness name: mmlu metrics:
  • —name: accnorm type: accnorm value: 0.2822 verified: false
  • —task: type: text-generation dataset: type: lm-evaluation-harness name: musr metrics:
  • —name: accnorm type: accnorm value: 0.3648 verified: false
  • —task: type: text-generation dataset: type: lm-evaluation-harness name: hellaswag metrics:
  • —name: acc type: acc value: 0.5141 verified: false
  • —name: accnorm type: accnorm value: 0.6793 verified: false

image/png

Eval

The fine tuned model (DevQuasar/analyticalreasoningr16a32_unsloth-Llama-3.2-3B-Instruct-bnb-4bit) has gained performace over the base model (unsloth/Llama-3.2-3B-Instruct-bnb-4bit) in the following tasks.

TestBase ModelFine-Tuned ModelPerformance Gain
leaderboardbbhlogicaldeductionseven_objects0.25200.43600.1840
leaderboardbbhlogicaldeductionfive_objects0.35600.45600.1000
leaderboardmusrteam_allocation0.22000.32000.1000
leaderboardbbhdisambiguation_qa0.30400.37600.0720
leaderboardgpqadiamond0.22220.27270.0505
leaderboardbbhmovie_recommendation0.59600.63600.0400
leaderboardbbhformal_fallacies0.50800.54000.0320
leaderboardbbhtrackingshuffledobjectsthreeobjects0.31600.34400.0280
leaderboardbbhcausal_judgement0.54550.56680.0214
leaderboardbbhweboflies0.49600.51600.0200
leaderboardmathgeometry_hard0.04550.06060.0152
leaderboardmathnumtheoryhard0.05190.06490.0130
leaderboardmusrmurder_mysteries0.52800.54000.0120
leaderboardgpqaextended0.27110.28020.0092
leaderboardbbhsports_understanding0.59600.60400.0080
leaderboardmathintermediatealgebrahard0.01070.01430.0036

Framework versions

  • —unsloth 2024.11.5
  • —trl 0.12.0

Training HW

  • —V100

I'm doing this to 'Make knowledge free for everyone', using my personal time and resources.

If you want to support my efforts please visit my ko-fi page: https://ko-fi.com/devquasar

Also feel free to visit my website https://devquasar.com/