norwegian
norwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.NorwegianCourtsBitextMining
NorwegianCourtsBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
Nynorsk and Bokmål parallel corpus from Norwegian courts. Norwegian courts have two standardised written languages. Bokmål is a variant closer to Danish, while Nynorsk was created to resemble regional dialects of Norwegian.
Task category
t2t
Domains
Legal, Written
Reference
https://opus.nlpl.eu/index.php
How to evaluate on this task
You can evaluate an embedding model on this… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NorwegianCourtsBitextMining.norwegian-courts
Norwegian Courts
Parallel corpus of Nynorsk and Bokmål from Norwegian Court transcriptions.
The data originates from the OPUS project.
Norwegian_idioms
NorEval: NorIdiom
This dataset is a part of the NorEval evaluation suite.See the NorEval codebase here: https://github.com/ltgoslo/norevalRead the preprint here: https://arxiv.org/abs/2504.07749
@article{mikhailov2025noreval,
title={NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark},
author={Mikhailov, Vladislav and Enstad, Tita and Samuel, David and Farseth{\aa}s, Hans Christian and Kutuzov, Andrey and Velldal, Erik and {\O}vrelid, Lilja}… See the full description on the dataset page: https://huggingface.co/datasets/Sprakbanken/Norwegian_idioms.norwegian_parliament
Dataset Card Creation Guide
Dataset Summary
This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.flan-norwegian
