CoolFace
Modelpublic

RichardErkhov/mpasila_-_Llama-3.2-Finnish-Wikipedia-1B-gguf

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes440downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

Llama-3.2-Finnish-Wikipedia-1B - GGUF

  • —Model creator: https://huggingface.co/mpasila/
  • —Original model: https://huggingface.co/mpasila/Llama-3.2-Finnish-Wikipedia-1B/

Original model description: --- base_model: unsloth/Llama-3.2-1B language:

  • —en
  • —fi license: llama3.2 tags:
  • —text-generation-inference
  • —transformers
  • —unsloth
  • —llama
  • —trl
  • —sft datasets:
  • —wikimedia/wikipedia --- Here's a "continued pre-trained" model using Finnish Wikipedia dataset. I still don't understand why no one in Finland has figured out that they could just do continued pre-training on existing models that are already supported by every frontend.. I've seen Japanese models perform pretty well with that kind of continued pre-training, yet Finnish models are still done from scratch which means they suck ass. If you compare them to Llama 3 or Gemma 2 they just suck so much. They can't even match Mistral 7B a model from last year. Just stop wasting money on training models from scratch, use these better models as base and train it on all your closed-source data I don't have access to. Thank you.

LoRA: mpasila/Llama-3.2-Finnish-Wikipedia-LoRA-1B

Trained with regular LoRA (not quantized/QLoRA) and LoRA rank was 128 and Alpha set to 32. Trained for 1 epoch using RTX 4090 for about 12,5 hours.

So it does have some issues but I could try training it on Gemma 2 2B and see if that's a better model for this (Gemma 2 already is better at Finnish than Llama 3) and maybe add more datasets containing Finnish.

Evaluation

ModelSizeTypeFIN-bench (score)Without math
mpasila/Llama-3.2-Finnish-Wikipedia-1B1BBase0.31700.4062
unsloth/Llama-3.2-1B1BBase0.40290.3881
Finnish-NLP/llama-7b-finnish7BBase0.23500.4203
LumiOpen/Viking-7B (1000B)7BBase0.37210.4453
HPLT/gpt-7b-nordic-prerelease7BBase0.31690.4524

Source

FIN-bench scores:
TaskVersionMetricValueStderr
bigbench_analogies0multiplechoicegrade0.4846±0.0440
bigbencharithmetic1digitaddition0multiplechoicegrade0.0300±0.0171
bigbencharithmetic1digitdivision0multiplechoicegrade0.0435±0.0435
bigbencharithmetic1digitmultiplication0multiplechoicegrade0.0200±0.0141
bigbencharithmetic1digitsubtraction0multiplechoicegrade0.0700±0.0256
bigbencharithmetic2digitaddition0multiplechoicegrade0.2200±0.0416
bigbencharithmetic2digitdivision0multiplechoicegrade0.0800±0.0273
bigbencharithmetic2digitmultiplication0multiplechoicegrade0.2400±0.0429
bigbencharithmetic2digitsubtraction0multiplechoicegrade0.1800±0.0386
bigbencharithmetic3digitaddition0multiplechoicegrade0.3300±0.0473
bigbencharithmetic3digitdivision0multiplechoicegrade0.2100±0.0409
bigbencharithmetic3digitmultiplication0multiplechoicegrade0.3000±0.0461
bigbencharithmetic3digitsubtraction0multiplechoicegrade0.5500±0.0500
bigbencharithmetic4digitaddition0multiplechoicegrade0.2800±0.0451
bigbencharithmetic4digitdivision0multiplechoicegrade0.2500±0.0435
bigbencharithmetic4digitmultiplication0multiplechoicegrade0.1500±0.0359
bigbencharithmetic4digitsubtraction0multiplechoicegrade0.4400±0.0499
bigbencharithmetic5digitaddition0multiplechoicegrade0.5100±0.0502
bigbencharithmetic5digitdivision0multiplechoicegrade0.3000±0.0461
bigbencharithmetic5digitmultiplication0multiplechoicegrade0.3100±0.0465
bigbencharithmetic5digitsubtraction0multiplechoicegrade0.4000±0.0492
bigbenchcauseandeffectone_sentence0multiplechoicegrade0.5882±0.0696
bigbenchcauseandeffectonesentenceno_prompt0multiplechoicegrade0.3922±0.0690
bigbenchcauseandeffecttwo_sentences0multiplechoicegrade0.4510±0.0704
bigbench_emotions0multiplechoicegrade0.1938±0.0313
bigbenchempiricaljudgments0multiplechoicegrade0.3434±0.0480
bigbenchgeneralknowledge0multiplechoicegrade0.2714±0.0535
bigbenchhhhalignment_harmless0multiplechoicegrade0.3966±0.0648
bigbenchhhhalignment_helpful0multiplechoicegrade0.3729±0.0635
bigbenchhhhalignment_honest0multiplechoicegrade0.3390±0.0622
bigbenchhhhalignment_other0multiplechoicegrade0.5581±0.0766
bigbenchintentrecognition0multiplechoicegrade0.0925±0.0110
bigbench_misconceptions0multiplechoicegrade0.4403±0.0430
bigbench_paraphrase0multiplechoicegrade0.5000±0.0354
bigbenchsentenceambiguity0multiplechoicegrade0.4833±0.0651
bigbenchsimilaritiesabstraction0multiplechoicegrade0.5921±0.0567

Uploaded Llama-3.2-Finnish-Wikipedia-1B model

  • —Developed by: mpasila
  • —License: Llama 3.2 Community License Agreement
  • —Finetuned from model : unsloth/Llama-3.2-1B

This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>