datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
norwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.NorwegianCourtsBitextMining
NorwegianCourtsBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
Nynorsk and Bokmål parallel corpus from Norwegian courts. Norwegian courts have two standardised written languages. Bokmål is a variant closer to Danish, while Nynorsk was created to resemble regional dialects of Norwegian.
Task category
t2t
Domains
Legal, Written
Reference
https://opus.nlpl.eu/index.php
How to evaluate on this task
You can evaluate an embedding model on this… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NorwegianCourtsBitextMining.norwegian-courts
Norwegian Courts
Parallel corpus of Nynorsk and Bokmål from Norwegian Court transcriptions.
The data originates from the OPUS project.
Norwegian_idioms
NorEval: NorIdiom
This dataset is a part of the NorEval evaluation suite.See the NorEval codebase here: https://github.com/ltgoslo/norevalRead the preprint here: https://arxiv.org/abs/2504.07749
@article{mikhailov2025noreval,
title={NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark},
author={Mikhailov, Vladislav and Enstad, Tita and Samuel, David and Farseth{\aa}s, Hans Christian and Kutuzov, Andrey and Velldal, Erik and {\O}vrelid, Lilja}… See the full description on the dataset page: https://huggingface.co/datasets/Sprakbanken/Norwegian_idioms.norwegian_parliament
Dataset Card Creation Guide
Dataset Summary
This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.norwegian-dyna-instruct
🧨 Norwegian dyna-instruct
Version
0.1.0 (changelog)
Languages
Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input
License
Mixed open licenses; see the table below
Sources
Five datasets (source cards)
Dataset Description
Number of samples: 14.40K
Number of tokens (Llama 3): 6.27M
Average conversation length in tokens (min, max): 435.63 (4, 8.92K)
Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.flan-norwegianwiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.norwegian_parliamentms-marco-norwegian
MS MARCO Norwegian
Norwegian translation of MS MARCO — 8,841,823 passages and 808,731 queries — for training and evaluating Norwegian retrieval and ranking models.
Translation
Bokmål (corpus): translated from English with TranslateGemma 12B (FP8) served via vLLM. The English source is the passage-ranking distribution of Microsoft's MS MARCO, as redistributed in HF Parquet form by sentence-transformers/msmarco.
Nynorsk (corpus_nn): translated from the Bokmål corpus… See the full description on the dataset page: https://huggingface.co/datasets/thivy/ms-marco-norwegian.norwegian-100h-v2norwegian-alpaca
NB Alpaca Norwegian Bokmål
This dataset is a translation to Norwegian Bokmål of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford.
An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI.
NorwegianParliamentClassification
NorwegianParliamentClassification
An MTEB dataset
Massive Text Embedding Benchmark
Norwegian parliament speeches annotated for sentiment
Task category
t2c
Domains
Government, Spoken
Reference
https://huggingface.co/datasets/NbAiLab/norwegian_parliament
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["NorwegianParliamentClassification"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NorwegianParliamentClassification.goldfish-Dp-norwegian-100mbgoldfish-Dp-norwegian-10mbreasoning_norwegian
Norwegian Reasoning
A reasoning dataset made by DeepSeek R1. The reasoning data is made from punctuation-restoration tasks from Wikipedia. We have stored the reasoning in cases where the output is 100% true.
A total of 22.000 tasks where generated.
Of these a total of 7794 tasks had the correct answer and where in Norwegian. This were trimmed to 6745 to be of the same size as the English reasoning dataset.
This was split into test=250, validation=250 and train=6245
norwegian-100h-v3reasoning_chat_norwegiannorwegian-xsum-nob
XSUM - Translated Norwegian Bokmål
Sourced from https://huggingface.co/datasets/NbAiLab/norwegian-xsum. Loaded from provided gzips and reuploaded due to errors accessing the original dataset through the dataset apis.
norwegian-paws-xNorwegian PAWS-X, Bokmaal and Nynorsk machine-translated versions of PAWS-X.
PAWS-X, a multilingual version of PAWS (Paraphrase Adversaries from Word Scrambling) for six languages.
This dataset contains 23,659 human translated PAWS evaluation pairs and 296,406 machine
translated training pairs in six typologically distinct languages: French, Spanish, German,
Chinese, Japanese, and Korean. English language is available by default. All translated
pairs are sourced from examples in PAWS-Wiki.
For further details, see the accompanying paper: PAWS-X: A Cross-lingual Adversarial Dataset
for Paraphrase Identification (https://arxiv.org/abs/1908.11828)
NOTE: There might be some missing or wrong labels in the dataset and we have replaced them with -1.Norwegian-Synthetic-HR-data-v-1
Synthetic norwegian public sector HR dataset
Dataset description
This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector.
The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use.
The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.alpaca_norwegian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_norwegian_taco.synthetic-from-classification-tasks-norwegian
Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset
The purpose of this dataset is to pre- or post-train embedding models for classification tasks.
The dataset consists of 100,000 samples generated with gemma-2-27b-it.
The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output.
Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-classification-tasks-norwegian.norwegian-ner-combined
Norwegian NER Combined Dataset
Dataset Description
This dataset combines the NorNE (Norwegian Named Entities) and WikiANN Norwegian datasets for Named Entity Recognition (NER) in Norwegian (Bokmål and Nynorsk).
Key Features
✅ 49,870 training samples (NorNE + WikiANN combined)
✅ 14,289 validation samples
✅ 13,450 test samples
✅ 4 entity types: PER, ORG, LOC, MISC
✅ Quality filtered: 12 problematic samples removed from NorNE
✅ Entity remapping: 9 original types… See the full description on the dataset page: https://huggingface.co/datasets/thivy/norwegian-ner-combined.norwegian-parliament-speeches
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Speeches from the Norwegian parliament from 1998 and 2022. Parsed from the Norwegian part of the EU ParlaMint, ParlaMint-NO
Dataset Sources
Source: https://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-77/
synthetic-from-text-mathing-short-tasks-norwegian
Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset
The purpose of this dataset is to pre- or post-train embedding models for text matching tasks on short texts.
The dataset consists of 100,000 samples generated with gemma-2-27b-it.
The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output.
Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-text-mathing-short-tasks-norwegian.Norwegian_sentimentalpaca-norwegian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-norwegian-cleaned.nor_wiki_reasoning_norwegiannorwegian_nynorsk_no
Norwegian Nynorsk Bible (1921)
Description
The Studentmållagsbibelen (Student Language Society Bible) of 1921 is the first complete Bible translation into Norwegian Nynorsk (New Norwegian), the written standard based on rural Norwegian dialects. Prepared by the Studentmållaget (Student Language Society) in Oslo, this translation from the original Hebrew and Greek was a landmark for the Nynorsk language movement. It includes the Protestant canon (66 books).… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/norwegian_nynorsk_no.
