datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Norwegian-Speech-DatasetYodaLingua-Norwegian
YodaLingua-Norwegian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Norwegian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
33,543 audio–transcription pairs
Total duration
93 hours
Speakers
1,608 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Norwegian.
