datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.reasoning_norwegian
Norwegian Reasoning
A reasoning dataset made by DeepSeek R1. The reasoning data is made from punctuation-restoration tasks from Wikipedia. We have stored the reasoning in cases where the output is 100% true.
A total of 22.000 tasks where generated.
Of these a total of 7794 tasks had the correct answer and where in Norwegian. This were trimmed to 6745 to be of the same size as the English reasoning dataset.
This was split into test=250, validation=250 and train=6245
reasoning_chat_norwegiannorwegian_nynorsk_no
Norwegian Nynorsk Bible (1921)
Description
The Studentmållagsbibelen (Student Language Society Bible) of 1921 is the first complete Bible translation into Norwegian Nynorsk (New Norwegian), the written standard based on rural Norwegian dialects. Prepared by the Studentmållaget (Student Language Society) in Oslo, this translation from the original Hebrew and Greek was a landmark for the Nynorsk language movement. It includes the Protestant canon (66 books).… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/norwegian_nynorsk_no.norwegian_bible_1930_no
Norwegian Bible (1930)
Description
The Norwegian Bible translation of 1930 (Det Norsk Bibelselskap 1930) is a revision of the older Norwegian Bible, prepared by the Norwegian Bible Society. It was translated from the original Hebrew and Greek texts and became the standard Norwegian Bible for much of the 20th century. It includes the Protestant canon (66 books).
Dataset Structure
Each row represents one Bible verse.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/norwegian_bible_1930_no.norwegian_nynorsk_1921_no
Norsk Studentmållagsbibelen (1921)
Description
The Studentmållagsbibelen — a Norwegian Nynorsk translation published in 1921 by the Norwegian Student Union Language Society (Studentmållaget). This is a landmark translation in the Nynorsk language, representing the first complete Bible in Norway's second official written standard.
Source: CrossWire Bible Society electronic text.
License: Public Domain
Dataset Structure
Each row represents one Bible… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/norwegian_nynorsk_1921_no.sst2-norwegian-bokmaal
Norwegian Translated SST-2 Dataset
Dataset
Overview
The dataset is a Norwegian machine-translation of the Stanford Sentiment Treebank (SST-2). The original dataset comprises sentences extracted from movie reviews, accompanied by human annotations indicating their sentiment.
Dataset Structure
The dataset has the following structure:
{
"idx": int,
"sentence": str,
"label": int,
"sentence_nob": str
}
Data Fields
idx:… See the full description on the dataset page: https://huggingface.co/datasets/Kushtrim/sst2-norwegian-bokmaal.
