samromur
Datasets
All datasets matching “samromur”samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.samromur_milljonSamrómur Milljón consists of approximately 1 million of speech recordings (967 hours) collected through the platform samromur.is; the transcripts accompanying these recordings were automatically verified using various ASR systems such as: Wav2Vec, Whisper and NeMo.samromur_asrSamrómur Icelandic Speech 1.0.samromur-500h-test
samromur-500h (test splits only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the test-related
splits of palli23/samromur-500h -- the training pool is not
redistributed here. Drawn from the Samrómur Milljón corpus.
test: the full official test split (9,308 utt., 8.02h).
test_4h: first-in-order subset of… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-500h-test.samromur-21.05-test
samromur-21.05 (test splits only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the test-related
splits of palli23/samromur-21.05 -- the training pool is not
redistributed here.
test: the official test split (10000 utt., 15.9h, 34% OOV
against the paper's independently-trained 21.05 scaling set).… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-21.05-test.samromur-milljon-test
samromur-milljon (test split only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the official test
split of palli23/samromur-milljon-splits (5017 utt.) -- the training
pool is not redistributed here. Used in the paper as the "Miljon, raw"
zero-shot difficulty anchor (92.5% of its prompts share exact text with… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-milljon-test.
