datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
norwegian-100h-v2norwegian-100h-v3norwegian-100hnrk-norwegian-speech-sample-v1
NRK Norwegian Speech Dataset (Sample)
Dataset Description
Note: This is a sample dataset containing a subset of chunks for demonstration and preview purposes.
The full dataset is available privately.
This dataset contains Norwegian speech data from NRK TV sports broadcasts, processed for automatic speech recognition (ASR) evaluation and research.
Dataset Statistics
Total chunks: 3
Episodes: 1
Total duration: 0.01 hours
Chunk types: subtitle_aligned… See the full description on the dataset page: https://huggingface.co/datasets/NRK-KIHUB/nrk-norwegian-speech-sample-v1.Norwegian-Speech-Datasetcommon-voice-norwegiannorwegian-nynorsk-speech-datasetYodaLingua-Norwegian
YodaLingua-Norwegian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Norwegian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
33,543 audio–transcription pairs
Total duration
93 hours
Speakers
1,608 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Norwegian.norwegian-bokm-l-speech-dataset
