CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Sprakbanken /Norwegian_idioms NorEval: NorIdiom This dataset is a part of the NorEval evaluation suite.See the NorEval codebase here: https://github.com/ltgoslo/norevalRead the preprint here: https://arxiv.org/abs/2504.07749 @article{mikhailov2025noreval, title={NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark}, author={Mikhailov, Vladislav and Enstad, Tita and Samuel, David and Farseth{\aa}s, Hans Christian and Kutuzov, Andrey and Velldal, Erik and {\O}vrelid, Lilja}… See the full description on the dataset page: https://huggingface.co/datasets/Sprakbanken/Norwegian_idioms.texttext-generation1K<n<10K6 likes408 downloads1y agoHugging Face02NbAiLab /norwegian_parliament Dataset Card Creation Guide Dataset Summary This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.texttext-classification1K<n<10K5 likes329 downloads2y agoHugging Face03pere /wiki_paragraphs_norwegian WIKI Paragraphs Norwegian A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.tabulartext-generation1M<n<10M0 likes138 downloads2y agoHugging Face04pere /reasoning_norwegian Norwegian Reasoning A reasoning dataset made by DeepSeek R1. The reasoning data is made from punctuation-restoration tasks from Wikipedia. We have stored the reasoning in cases where the output is 100% true. A total of 22.000 tasks where generated. Of these a total of 7794 tasks had the correct answer and where in Norwegian. This were trimmed to 6745 to be of the same size as the English reasoning dataset. This was split into test=250, validation=250 and train=6245 tabulartext-generation1K<n<10K1 likes52 downloads2y agoHugging Face05pere /reasoning_chat_norwegiantabular1K<n<10K1 likes37 downloads2y agoHugging Face06Hebbelille /Norwegian-Synthetic-HR-data-v-1 Synthetic norwegian public sector HR dataset Dataset description This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector. The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use. The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.texttext-generation1K<n<10K0 likes24 downloads10mo agoHugging Face07pere /nor_wiki_reasoning_norwegiantext1K<n<10K0 likes19 downloads2y agoHugging Face08buzzcraft /Norwegian_PII Norwegian PII (Placeholder) This dataset contains annotated Norwegian text samples where sensitive entities are replaced with placeholders such as name_1, address_1, email_1, and phone_1. It is intended as a safe-to-share placeholder version of a Norwegian PII extraction dataset. Dataset Summary The dataset preserves the original annotation structure while removing direct values for selected entity types. Placeholder tokens are used for: person address email… See the full description on the dataset page: https://huggingface.co/datasets/buzzcraft/Norwegian_PII.texttoken-classification1K<n<10K0 likes17 downloads4mo agoHugging Face09TokenHaven /FineWeb-Edu-Norwegian High Quality Norwegian Corpus This dataset contains a large collection of high-quality Norwegian text data with their metadata. To access the full data please visit Token Haven Creation The dataset was created by filtering all English common crawl data for high-quality text using the FineWeb-Edu classifier with education score of 4 or higher over 5. The data is source from the v1.0.0 of the HuggingFaceFW/fineweb-edu dataset which corresponds to CC-MAIN-2024-10… See the full description on the dataset page: https://huggingface.co/datasets/TokenHaven/FineWeb-Edu-Norwegian.texttext-generationn<1K0 likes7 downloads1y agoHugging Face10schneiderkamplab /dfm12-norwegian-inclusive-nb-samtale-pairs dfm12-norwegian-inclusive-nb-samtale-pairs Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-norwegian-inclusive-nb-samtale-pairs.texttext-generationn<1K0 likes5h agoHugging Face11schneiderkamplab /dfm12-norwegian-inclusive-reasoning-norwegian dfm12-norwegian-inclusive-reasoning-norwegian Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-norwegian-inclusive-reasoning-norwegian.texttext-generation1K<n<10K0 likes5h agoHugging Face12schneiderkamplab /dfm12-norwegian-magpie-qwen3-bokmaal dfm12-norwegian-magpie-qwen3-bokmaal Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-norwegian-magpie-qwen3-bokmaal.texttext-generationn<1K0 likes5h agoHugging Face13schneiderkamplab /dfm12-norwegian-nb-samtale-pairs dfm12-norwegian-nb-samtale-pairs Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-norwegian-nb-samtale-pairs.texttext-generation1K<n<10K0 likes5h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.