LokaalHub/sv-SE-asr-cv
Swedish ASR (Common Voice 22, filtered + rebalanced) Swedish (sv-SE) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open fsicoli/common_voice_22_0 mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train 22166 26.1 dev 694 0.8 test 1602 2.0 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/sv-SE-asr-cv.
Swedish ASR (Common Voice 22, filtered + rebalanced)
Swedish (sv-SE) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open `fsicoli/common_voice_22_0` mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
train = the official validated train split + the filtered other bucket + the excess `dev`/`test` speakers: Common Voice's official dev/test are balanced for benchmarking and far too large to hold out when fine-tuning a low-resource language, so they are capped (dev ~0.75h, test ~2.0h) by whole speakers and the remainder moved into train. Held-out dev/test stay disjoint from train by both speaker and sentence.
Each unvalidated other clip is CTC-decoded with `KBLab/wav2vec2-large-voxrex-swedish` and kept only if its character error rate vs the prompt is low (CER <= 0.25); 5451/5564 kept.
Columns
audio (16 kHz mono), text, lang (sv-SE), client_id (anonymized speaker), split, source (validated | other | moved-validated), quality_score.
