CoolFace
Datasetpublic

NbAiLab/nb-asr-qwen3whisperxagreement-v1

nb-asr-qwen3whisperxagreement-v1 Word-level forced alignment training data for Norwegian speech, produced by keeping only examples where two independent aligners — WhisperX and Qwen3 (Lunde forced aligner) — agree within a tight tolerance. Dataset Description This dataset contains 702,067 speech segments drawn from the NB-ASR Norwegian audio corpus. Each record pairs an audio file with a word-level forced alignment in a format suitable for training a… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-qwen3whisperxagreement-v1.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes10downloads
settings

This repository belongs to NbAiLab on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenb-asr-qwen3whisperxagreement-v1
visibilitypublic
licencecc-by-4.0
gatedno
ownerNbAiLab
Account settings
NbAiLab/nb-asr-qwen3whisperxagreement-v1 · CoolFace