CoolFace
Modelpublic

HPLT/NorOLMo-13B

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
3likes9.4kdownloads
Model Card

NorOLMo

NorOLMo 1.0 is a base (not instruction-tuned) large language model, continually pre-trained on Norwegian and Scandinavian data starting from the English OLMo2-13B model.

It is the only modern fully open large language model specifically adapted for Norwegian and Sámi. At the same time, its performance is comparable to other multilingual open models (see the Evaluation section below). "Fully open" here means that the entirety of the trainining data and training pipeline is transparent and reproducible, not only that the model weights are available.

The model was trained for 33 000 steps on around 275 billion tokens. Maximum sequence length was set to 4 096 in the beginning of training, and then was extended to 16 384 starting from the checkpoint step_31000. Intermediate checkpoints are published here as branches. The main branch contains the model's weights after step 33 000 (stage 3).

In a few months, we hope to follow with post-trained versions of NorOLMo.

Evaluation

Below is a comparison of fully-open models supporting Norwegian. The figure shows the aggregate score across all 35 NorEval 1.1 tasks (5 categories, category average). Scores are first averaged within each task category, then averaged across categories. This gives equal weight to each category regardless of how many tasks it contains. Each task score is normalized to a 0–100 scale where 0 = random baseline performance and 100 = perfect score, then averaged across tasks. This accounts for different chance levels across tasks (e.g. 25% for 4-choice QA vs. 50% for binary classification).

More detailed evaluation is evailable in our interactive NorEval dashboard: https://ltgoslo.github.io/llm-dashboard.

[image]

Furthermore, an interactive per-checkpoint evaluation with additional ablation studies is available here.

Data Details

Stage 1 (24 000 steps -- 200B tokens)

Data ("pretraining data")

<details>

<summary>Data splits</summary>

DataPercentageUnique TokensTotal TokensNumber of DocumentsAverage Document Length
HPLT Bokmål39.5739.8B79.7B36.5M1 092
HPLT Nynorsk4.951.2B10.0B1.5M826
HPLT Faroese0.460.2B0.9B0.3M711
HPLT Icelandic2.505.0B5.0B4.3M1 173
HPLT Swedish12.0992.1B24.4B97.7M942
HPLT Danish12.1250.1B24.4B52.5M954
FinePDFs Bokmål8.368.4B16.8B1.5M5 604
FinePDFs Nynorsk1.150.3B2.3B92.8K3 117
FinePDFs Faroese0.1787.1M0.3B20.8K4 196
FinePDFs Icelandic1.603.2B3.2B0.4M8 855
FinePDFs Swedish2.4818.9B5.0B4.1M4 574
FinePDFs Danish2.4510.1B4.9B2.4M4 190
Northern Sami0.1846.4M0.4B0.2M288
Wiki (OLMo-Mix)0.020.2B40.3M0.3M667
Alg. Stack (OLMo-Mix)0.040.6B80.5M0.1M4 201
Open Web Math (OLMo-Mix)0.040.6B80.5M0.1M4 199
ArXiv (OLMo-Mix)0.051.0B0.1B0.2M5 210
PeS2o (OLMo-Mix)0.152.5B0.3B1.6M1 641
DCLM (OLMo-Mix)9.5048.3B19.1B35.1M1 377
StarCoder (OLMo-Mix)2.1030.5B4.2B23.6M1 293
[!NOTE] The number of documents represents the total unique number of documents, not the documents used during training.
[!NOTE] We only took a portion of OLMo-Mix as our unique data.

</details>

Stage 2 (6 000 steps -- 50B tokens) and Stage 3 (3 000 steps -- 25B tokens)

Data ("midtraining data")

<details>

<summary>Data splits</summary>

Data Splits | Data | Percentage | Unique Tokens | Total Tokens | Number of Documents | Average Document Length | | ------------------------ | ---------- | ------------- | ------------ | ------------------- | ----------------------- | | HPLT Bokmål | 45.78 | 23.0B | 23.0B | 19.0M | 1 215 | | HPLT Nynorsk | 7.84 | 1.0B | 3.9B | 1.0M | 1 003 | | HPLT Icelandic | 6.87 | 3.5B | 3.5B | 2.7M | 1 268 | | HPLT Swedish | 4.90 | 2.5B | 2.5B | 3.6M | 3 403 | | HPLT Danish | 7.73 | 3.9B | 3.9B | 4.1M | 2 950 | | FinePDFs-Edu Bokmål | 2.24 | 1.1B | 1.1B | 0.2M | 6 897 | | FinePDFs-Edu Nynorsk | 0.28 | 35.8M | 0.1B | 9.7K | 3 681 | | FinePDFs Faroese | 0.69 | 87.1M | 0.3B | 20.8K | 4 196 | | FinePDFs-Edu Icelandic | 0.53 | 0.3B | 0.3B | 40.1K | 6 598 | | FinePDFs-Edu Swedish | 5.80 | 2.9B | 2.9B | 0.4M | 6 755 | | FinePDFs-Edu Danish | 2.97 | 1.5B | 1.5B | 0.3M | 5 833 | | FinePDFs-Edu English | 7.00 | 7.2B | 3.5B | 1.1M | 6 280 | | Northern Sami | 0.37 | 46.4M | 0.2B | 0.2M | 288 | | Stack-Edu | 5.00 | 12.8B | 2.5B | 15.0M | 856 | | MegaMath Web-Pro | 0.84 | 13.7B | 0.4B | 15.0M | 917 | | FineMath 4+ | 0.62 | 10.1B | 0.3B | 6.7M | 1 512 | | InfiWebMath 4+ | 0.54 | 8.9B | 0.3B | 6.3M | 1 417 |

</details>

Training details

<details>

<summary>Stage 1</summary>

HyperparameterValue
Embedding train steps1 000
Warmup steps2 000
Total train steps24 000
Learning rate scheduleWarmup + constant
Learning rate3e-4
Weight decay1e-1
Sequence length4 096
Batch size2 048
RoPe theta500 000
Clip grad1.0
Adam epsilon1e-8
Adam beta_10.9
Adam beta_20.95
RMSNorm epsilon1e-6
Z-loss ratio1e-5
Diffusion loss ratio2e-2

</details>

<details>

<summary>Stage 2</summary>

HyperparameterValue
Decay steps6 000
Total train steps6 000
Learning rate scheduleLinear decay
Initial learning rate3e-4
Final learning rate1.5e-4
Weight decay1e-1
Sequence length4 096
Batch size2 048
RoPe theta500 000
Clip grad1.0
Adam epsilon1e-8
Adam beta_10.9
Adam beta_20.95
RMSNorm epsilon1e-6
Z-loss ratio1e-5
Diffusion loss ratio2e-2

</details>

<details>

<summary>Stage 3</summary>

HyperparameterValue
Decay steps3 000
Total train steps3 000
Learning rate scheduleLinear decay
Max learning rate1.5e-4
Final learning rate0
Weight decay1e-1
Sequence length16 384
Batch size512
RoPe theta2 000 000
Clip grad1.0
Adam epsilon1e-8
Adam beta_10.9
Adam beta_20.95
RMSNorm epsilon1e-6
Z-loss ratio1e-5
Diffusion loss ratio2e-2

</details>

Acknowledgements

Training was conducted as a part of the HPLT project.

This project has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No 101070350 and from UK Research and Innovation (UKRI) under the UK government’s Horizon Europe funding guarantee [grant number 10052546]