CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes2k downloads18d agoHugging Face02mteb /NorwegianCourtsBitextMining NorwegianCourtsBitextMining An MTEB dataset Massive Text Embedding Benchmark Nynorsk and Bokmål parallel corpus from Norwegian courts. Norwegian courts have two standardised written languages. Bokmål is a variant closer to Danish, while Nynorsk was created to resemble regional dialects of Norwegian. Task category t2t Domains Legal, Written Reference https://opus.nlpl.eu/index.php How to evaluate on this task You can evaluate an embedding model on this… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NorwegianCourtsBitextMining.texttranslation1K<n<10K0 likes1.4k downloads1y agoHugging Face03kardosdrur /norwegian-courts Norwegian Courts Parallel corpus of Nynorsk and Bokmål from Norwegian Court transcriptions. The data originates from the OPUS project. textsentence-similarity1K<n<10K1 likes1.2k downloads3y agoHugging Face04Sprakbanken /Norwegian_idioms NorEval: NorIdiom This dataset is a part of the NorEval evaluation suite.See the NorEval codebase here: https://github.com/ltgoslo/norevalRead the preprint here: https://arxiv.org/abs/2504.07749 @article{mikhailov2025noreval, title={NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark}, author={Mikhailov, Vladislav and Enstad, Tita and Samuel, David and Farseth{\aa}s, Hans Christian and Kutuzov, Andrey and Velldal, Erik and {\O}vrelid, Lilja}… See the full description on the dataset page: https://huggingface.co/datasets/Sprakbanken/Norwegian_idioms.texttext-generation1K<n<10K6 likes408 downloads1y agoHugging Face05NbAiLab /norwegian_parliament Dataset Card Creation Guide Dataset Summary This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.texttext-classification1K<n<10K5 likes329 downloads2y agoHugging Face06danish-foundation-models /norwegian-dyna-instruct 🧨 Norwegian dyna-instruct Version 0.1.0 (changelog) Languages Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input License Mixed open licenses; see the table below Sources Five datasets (source cards) Dataset Description Number of samples: 14.40K Number of tokens (Llama 3): 6.27M Average conversation length in tokens (min, max): 435.63 (4, 8.92K) Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.imagequestion-answering10K<n<100K0 likes250 downloads18d agoHugging Face07Lauler /flan-norwegiantext1M<n<10M0 likes179 downloads2y agoHugging Face08pere /wiki_paragraphs_norwegian WIKI Paragraphs Norwegian A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.tabulartext-generation1M<n<10M0 likes138 downloads2y agoHugging Face09mteb /norwegian_parliamenttext1K<n<10K0 likes108 downloads1y agoHugging Face10thivy /ms-marco-norwegian MS MARCO Norwegian Norwegian translation of MS MARCO — 8,841,823 passages and 808,731 queries — for training and evaluating Norwegian retrieval and ranking models. Translation Bokmål (corpus): translated from English with TranslateGemma 12B (FP8) served via vLLM. The English source is the passage-ranking distribution of Microsoft's MS MARCO, as redistributed in HF Parquet form by sentence-transformers/msmarco. Nynorsk (corpus_nn): translated from the Bokmål corpus… See the full description on the dataset page: https://huggingface.co/datasets/thivy/ms-marco-norwegian.texttext-retrieval10M<n<100M1 likes96 downloads4mo agoHugging Face11davidilag /norwegian-100h-v2audio10K<n<100K0 likes72 downloads2y agoHugging Face12NbAiLab /norwegian-alpaca NB Alpaca Norwegian Bokmål This dataset is a translation to Norwegian Bokmål of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford. An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI. texttext-generation10K<n<100K10 likes65 downloads3y agoHugging Face13mteb /NorwegianParliamentClassification NorwegianParliamentClassification An MTEB dataset Massive Text Embedding Benchmark Norwegian parliament speeches annotated for sentiment Task category t2c Domains Government, Spoken Reference https://huggingface.co/datasets/NbAiLab/norwegian_parliament How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["NorwegianParliamentClassification"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/NorwegianParliamentClassification.texttext-classification1K<n<10K0 likes62 downloads1y agoHugging Face14fpadovani /goldfish-Dp-norwegian-100mbtext10M<n<100M0 likes59 downloads19d agoHugging Face15fpadovani /goldfish-Dp-norwegian-10mbtext10M<n<100M0 likes58 downloads19d agoHugging Face16pere /reasoning_norwegian Norwegian Reasoning A reasoning dataset made by DeepSeek R1. The reasoning data is made from punctuation-restoration tasks from Wikipedia. We have stored the reasoning in cases where the output is 100% true. A total of 22.000 tasks where generated. Of these a total of 7794 tasks had the correct answer and where in Norwegian. This were trimmed to 6745 to be of the same size as the English reasoning dataset. This was split into test=250, validation=250 and train=6245 tabulartext-generation1K<n<10K1 likes52 downloads2y agoHugging Face17davidilag /norwegian-100h-v3audio10K<n<100K0 likes49 downloads2y agoHugging Face18pere /reasoning_chat_norwegiantabular1K<n<10K1 likes37 downloads2y agoHugging Face19tollefj /norwegian-xsum-nob XSUM - Translated Norwegian Bokmål Sourced from https://huggingface.co/datasets/NbAiLab/norwegian-xsum. Loaded from provided gzips and reuploaded due to errors accessing the original dataset through the dataset apis. textsummarization100K<n<1M1 likes34 downloads3y agoHugging Face20NbAiLab /norwegian-paws-xNorwegian PAWS-X, Bokmaal and Nynorsk machine-translated versions of PAWS-X. PAWS-X, a multilingual version of PAWS (Paraphrase Adversaries from Word Scrambling) for six languages. This dataset contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean. English language is available by default. All translated pairs are sourced from examples in PAWS-Wiki. For further details, see the accompanying paper: PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification (https://arxiv.org/abs/1908.11828) NOTE: There might be some missing or wrong labels in the dataset and we have replaced them with -1.texttext-classification100K<n<1M2 likes30 downloads3y agoHugging Face21Hebbelille /Norwegian-Synthetic-HR-data-v-1 Synthetic norwegian public sector HR dataset Dataset description This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector. The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use. The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.texttext-generation1K<n<10K0 likes24 downloads10mo agoHugging Face22saillab /alpaca_norwegian_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_norwegian_taco.text10K<n<100K1 likes23 downloads2y agoHugging Face23ThatsGroes /synthetic-from-classification-tasks-norwegian Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset The purpose of this dataset is to pre- or post-train embedding models for classification tasks. The dataset consists of 100,000 samples generated with gemma-2-27b-it. The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output. Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-classification-tasks-norwegian.texttext-classification10K<n<100K0 likes22 downloads2y agoHugging Face24thivy /norwegian-ner-combined Norwegian NER Combined Dataset Dataset Description This dataset combines the NorNE (Norwegian Named Entities) and WikiANN Norwegian datasets for Named Entity Recognition (NER) in Norwegian (Bokmål and Nynorsk). Key Features ✅ 49,870 training samples (NorNE + WikiANN combined) ✅ 14,289 validation samples ✅ 13,450 test samples ✅ 4 entity types: PER, ORG, LOC, MISC ✅ Quality filtered: 12 problematic samples removed from NorNE ✅ Entity remapping: 9 original types… See the full description on the dataset page: https://huggingface.co/datasets/thivy/norwegian-ner-combined.texttoken-classification10K<n<100K0 likes22 downloads9mo agoHugging Face25trymtv /norwegian-parliament-speeches Dataset Card for Dataset Name Dataset Details Dataset Description Speeches from the Norwegian parliament from 1998 and 2022. Parsed from the Norwegian part of the EU ParlaMint, ParlaMint-NO Dataset Sources Source: https://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-77/ texttext-classification100K<n<1M0 likes21 downloads3y agoHugging Face26ThatsGroes /synthetic-from-text-mathing-short-tasks-norwegian Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset The purpose of this dataset is to pre- or post-train embedding models for text matching tasks on short texts. The dataset consists of 100,000 samples generated with gemma-2-27b-it. The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output. Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-text-mathing-short-tasks-norwegian.text10K<n<100K0 likes21 downloads2y agoHugging Face27sepidmnorozy /Norwegian_sentimenttext1K<n<10K0 likes19 downloads4y agoHugging Face28saillab /alpaca-norwegian-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-norwegian-cleaned.text10K<n<100K0 likes19 downloads2y agoHugging Face29pere /nor_wiki_reasoning_norwegiantext1K<n<10K0 likes19 downloads2y agoHugging Face30k-mktr /norwegian_nynorsk_no Norwegian Nynorsk Bible (1921) Description The Studentmållagsbibelen (Student Language Society Bible) of 1921 is the first complete Bible translation into Norwegian Nynorsk (New Norwegian), the written standard based on rural Norwegian dialects. Prepared by the Studentmållaget (Student Language Society) in Oslo, this translation from the original Hebrew and Greek was a landmark for the Nynorsk language movement. It includes the Protestant canon (66 books).… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/norwegian_nynorsk_no.tabular10K<n<100K0 likes19 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.