datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NPSC_test
Dataset Card for NBAiLab/NPSC
The Norwegian Parliament Speech Corpus (NPSC) is a corpus for training a Norwegian ASR (Automatic Speech Recognition) models. The corpus is created by Språkbanken at the National Library in Norway.
NPSC is based on sound recording from meeting in the Norwegian Parliament. These talks are orthographically transcribed to either Norwegian Bokmål or Norwegian Nynorsk. In addition to the data actually included in this dataset, there is a significant amount… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/NPSC_test.nb-asr-eval-withwav-sorted
NB-ASR Eval with WAV Sorted
Hardest-first copy of NbAiLab/nb-asr-eval-withwav for targeted human cleanup.
Rows are intended to be sorted independently within each split by ASR/WER difficulty.
Audio paths and split metadata layout are preserved so downstream tools can switch
from the original repo to NbAiLab/nb-asr-eval-withwav-sorted without changing file lookup logic.
After scoring, each metadata row may include original_source_index, priority_rank,
asr_wer, asr_cer… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-eval-withwav-sorted.nb-librivox
📄 NB-LibriVox
A high-quality Norwegian text-to-speech (TTS) dataset derived from LibriVox public domain audiobooks. It includes audio clips with pseudo-aligned transcripts and punctuation, curated by the National Library of Norway for speech synthesis and ASR research.
📂 Dataset Overview
Field
Description
file_name
Audio file name in .wav format
id
Unique identifier for each utterance/sample
text
Transcript automatically generated using NB-Whisper Large… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-librivox.
