CoolFace
Modelpublic

DevQuasar/analytical_reasoning_r16a32_unsloth-Llama-3.2-3B-Instruct-bnb-4bit_adapter

sourceHugging Facellama3.2updated 2y agoView on Hugging Face
1likes
Model Card

<img src="https://raw.githubusercontent.com/csabakecskemeti/devquasar/main/dq_logo_black-transparent.png" width="200"/>

'Make knowledge free for everyone'

<a href='https://ko-fi.com/L4L416YX7C' target='_blank'><img height='36' style='border:0px;height:36px;' src='https://storage.ko-fi.com/cdn/kofi6.png?v=6' border='0' alt='Buy Me a Coffee at ko-fi.com' /></a>

image/png

Eval

The fine tuned model (DevQuasar/analyticalreasoningr16a32_unsloth-Llama-3.2-3B-Instruct-bnb-4bit) has gained performace over the base model (unsloth/Llama-3.2-3B-Instruct-bnb-4bit) in the following tasks.

TestBase ModelFine-Tuned ModelPerformance Gain
leaderboardbbhlogicaldeductionseven_objects0.25200.43600.1840
leaderboardbbhlogicaldeductionfive_objects0.35600.45600.1000
leaderboardmusrteam_allocation0.22000.32000.1000
leaderboardbbhdisambiguation_qa0.30400.37600.0720
leaderboardgpqadiamond0.22220.27270.0505
leaderboardbbhmovie_recommendation0.59600.63600.0400
leaderboardbbhformal_fallacies0.50800.54000.0320
leaderboardbbhtrackingshuffledobjectsthreeobjects0.31600.34400.0280
leaderboardbbhcausal_judgement0.54550.56680.0214
leaderboardbbhweboflies0.49600.51600.0200
leaderboardmathgeometry_hard0.04550.06060.0152
leaderboardmathnumtheoryhard0.05190.06490.0130
leaderboardmusrmurder_mysteries0.52800.54000.0120
leaderboardgpqaextended0.27110.28020.0092
leaderboardbbhsports_understanding0.59600.60400.0080
leaderboardmathintermediatealgebrahard0.01070.01430.0036

Framework versions

  • —unsloth 2024.11.5
  • —trl 0.12.0

Training HW

  • —V100