datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
science-theory-textbooksNBA_Games
NBA Full-Game Video Dataset
This dataset provides metadata, official statistics, and official play-by-play annotations for full-length NBA game videos available on YouTube. Instead of redistributing video files, we provide YouTube video IDs and URLs so users can download videos independently when their use case and local policies allow it.
The dataset links long-form basketball videos with structured NBA.com game data. Each retained game has a verified… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/NBA_Games.wild-science-theory-textbooksNCC
Dataset Card for NbAiLab/NCC
⚠️ Important Update (December 2024)
Previously, newspapers were a significant part of the Norwegian Colossal Corpus (NCC), particularly the newspapers distributed under the so called "Språkbank-avtalen". As of December 2024, at the request of media houses, we have ceased distributing newspapers under this agreement, including the "Norsk Aviskorpus." However, NCC still includes numerous newspapers that are released under more open… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/NCC.art-theory-textbooksnbnn_language_detection
Dataset Card for Bokmål-Nynorsk Language Detection (main_train_split)
Dataset Summary
This dataset is intended for language detection for Bokmål to Nynorsk and vice versa. It contains 800,000 sentence pairs, sourced from Språkbanken and pruned to avoid overlap with the NorBench dataset. The data comes from translations of news text from Norsk telegrambyrå (NTB), performed by Nynorsk pressekontor (NPK). In addition the dev and test set has 1000 entries.
Data… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nbnn_language_detection.norwegian_parliament
Dataset Card Creation Guide
Dataset Summary
This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.nb_distil_speech_noconcat_stortinget
Dataset Card for NbAiLab/nb_distil_speech_noconcat_stortinget
Dataset Summary
NbAiLab/nb_distil_speech_noconcat_stortinget is a curated subset of the Stortinget Speech Corpus (SSC), a large-scale Norwegian parliamentary speech dataset. This subset focuses on non-concatenated speech segments and includes automatic transcriptions generated using OpenAI's Whisper model. It is designed to facilitate the development and evaluation of Automatic Speech Recognition (ASR) systems… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb_distil_speech_noconcat_stortinget.nb_distil_speech_noconcat_nstlunde_nor_nob_reading_optimisedTest only - not for training.
First version - 0.1 of lunde_nor_nob_reading_optimised
This dataset does not contain any audio data.
Export Details
Train samples: 10040932
Validation samples: 0
Test samples: 0
Dataset created using search datasets:lunde_nor_nob_reading_optimised.
nba-box-scoresdiverse-svg-prompts
Diverse SVG Prompts
Diverse SVG Prompts is a public collection of 20,000 high-quality,
generated and filtered English briefs for SVG and vector-graphics generation.
It contains 18,000 general illustration prompts and 2,000 lettering prompts.
Schema
The dataset intentionally has only two columns:
prompt: the complete visual brief.
type_tags: a list of category, author-model, and processing tags.
Example:
{
"prompt": "A moonlit mechanical heron..."… See the full description on the dataset page: https://huggingface.co/datasets/Nbardy/diverse-svg-prompts.nba-home-court-2017-2026
NBA Home-Court Advantage 2017-2026: 11,777 Games
What this is
Every NBA regular-season game across 10 seasons (2016-17 through 2025-26): 11,777 games, with home and away team, final score, scoring margin, venue, and a neutral-site flag. Collected from ESPN's official sports API through the MrBridge ESPN MCP Server. No scraping; the data is public game results from the official feed.
Companion study: NBA Home-Court Advantage: 10 Seasons, 11,777 Games — home… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/nba-home-court-2017-2026.nynorsk_norm_200eval
Nynorsk Norm 200eval
nynorsk_norm_200eval is a high-quality, small-scale parallel corpus comprising 200 Norwegian Bokmål–Nynorsk sentence pairs collected from official sources and public institutions. Each example includes:
nb: Original sentence in Bokmål
nn_original: Original Nynorsk sentence (typically an official translation)
nn_alt_original: Original Nynorsk sentence (typically an official translation) - alt version
nn_husnorm: Sentence rewritten in Nynorsk following an… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nynorsk_norm_200eval.NbAiLab__nb-llama-3.1-8B-Instruct-details
Dataset Card for Evaluation run of NbAiLab/nb-llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model NbAiLab/nb-llama-3.1-8B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NbAiLab__nb-llama-3.1-8B-Instruct-details.nynorsk_dpo
Bokmål–Nynorsk DPO
bokmal_nynorsk_dpo is a dataset for Direct Preference Optimization (DPO) training, focusing on Bokmål–Nynorsk translation.Each example consists of a prompt in Norwegian Bokmål and two candidate translations in Nynorsk:
prompt: Input sentence in Bokmål
chosen: Preferred Nynorsk translation (higher quality, closer to target norm)
rejected: Less preferred Nynorsk translation
This format enables reinforcement learning from human preferences, where models learn… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nynorsk_dpo.ndla_npk_conversational_nb_to_nn_tags_balanced
Balanced version of NbAiLab/ndla_npk_conversational_nb_to_nn with tags.
The corpus consists of:
70.000 samples from NbAiLab/ndla_npk_conversational_nb_to_nn
15.000 samples with tags from NDLA
15.000 samples with tags from NPL
TOTAL: 100k samples
The coprus is mainly made for GRPO-training
freddy-testDette er et datasett som skal slettes.
ndla_npk_conversational_nb_to_nn_tags_allcapsnba-experiment-queuetext-to-sql-nbanbanb-asr-qwen3whisperxagreement-v1
nb-asr-qwen3whisperxagreement-v1
Word-level forced alignment training data for Norwegian speech, produced by keeping only examples where two independent aligners — WhisperX and Qwen3 (Lunde forced aligner) — agree within a tight tolerance.
Dataset Description
This dataset contains 702,067 speech segments drawn from the NB-ASR Norwegian audio corpus. Each record pairs an audio file with a word-level forced alignment in a format suitable for training a… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-qwen3whisperxagreement-v1.NbAiLab__nb-llama-3.1-8B-sft-details
Dataset Card for Evaluation run of NbAiLab/nb-llama-3.1-8B-sft
Dataset automatically created during the evaluation run of model NbAiLab/nb-llama-3.1-8B-sft
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NbAiLab__nb-llama-3.1-8B-sft-details.nbanba-classifierRecipes_json_vector_nbarbconenba
