datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
norwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.Norwegian_idioms
NorEval: NorIdiom
This dataset is a part of the NorEval evaluation suite.See the NorEval codebase here: https://github.com/ltgoslo/norevalRead the preprint here: https://arxiv.org/abs/2504.07749
@article{mikhailov2025noreval,
title={NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark},
author={Mikhailov, Vladislav and Enstad, Tita and Samuel, David and Farseth{\aa}s, Hans Christian and Kutuzov, Andrey and Velldal, Erik and {\O}vrelid, Lilja}… See the full description on the dataset page: https://huggingface.co/datasets/Sprakbanken/Norwegian_idioms.norwegian-dyna-instruct
🧨 Norwegian dyna-instruct
Version
0.1.0 (changelog)
Languages
Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input
License
Mixed open licenses; see the table below
Sources
Five datasets (source cards)
Dataset Description
Number of samples: 14.40K
Number of tokens (Llama 3): 6.27M
Average conversation length in tokens (min, max): 435.63 (4, 8.92K)
Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.wiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.norwegian-alpaca
NB Alpaca Norwegian Bokmål
This dataset is a translation to Norwegian Bokmål of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford.
An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI.
reasoning_norwegian
Norwegian Reasoning
A reasoning dataset made by DeepSeek R1. The reasoning data is made from punctuation-restoration tasks from Wikipedia. We have stored the reasoning in cases where the output is 100% true.
A total of 22.000 tasks where generated.
Of these a total of 7794 tasks had the correct answer and where in Norwegian. This were trimmed to 6745 to be of the same size as the English reasoning dataset.
This was split into test=250, validation=250 and train=6245
Norwegian-Synthetic-HR-data-v-1
Synthetic norwegian public sector HR dataset
Dataset description
This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector.
The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use.
The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.FineWeb-Edu-Norwegian
High Quality Norwegian Corpus
This dataset contains a large collection of high-quality Norwegian text data with their metadata.
To access the full data please visit Token Haven
Creation
The dataset was created by filtering all English common crawl data for high-quality text using the FineWeb-Edu classifier with education score of 4 or higher over 5.
The data is source from the v1.0.0 of the HuggingFaceFW/fineweb-edu dataset which corresponds to CC-MAIN-2024-10… See the full description on the dataset page: https://huggingface.co/datasets/TokenHaven/FineWeb-Edu-Norwegian.dfm12-norwegian-inclusive-nb-samtale-pairs
dfm12-norwegian-inclusive-nb-samtale-pairs
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-norwegian-inclusive-nb-samtale-pairs.dfm12-norwegian-inclusive-reasoning-norwegian
dfm12-norwegian-inclusive-reasoning-norwegian
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-norwegian-inclusive-reasoning-norwegian.dfm12-norwegian-magpie-qwen3-bokmaal
dfm12-norwegian-magpie-qwen3-bokmaal
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-norwegian-magpie-qwen3-bokmaal.dfm12-norwegian-nb-samtale-pairs
dfm12-norwegian-nb-samtale-pairs
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-norwegian-nb-samtale-pairs.
