datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.samromur-500h-test
samromur-500h (test splits only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the test-related
splits of palli23/samromur-500h -- the training pool is not
redistributed here. Drawn from the Samrómur Milljón corpus.
test: the full official test split (9,308 utt., 8.02h).
test_4h: first-in-order subset of… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-500h-test.samromur-21.05-test
samromur-21.05 (test splits only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the test-related
splits of palli23/samromur-21.05 -- the training pool is not
redistributed here.
test: the official test split (10000 utt., 15.9h, 34% OOV
against the paper's independently-trained 21.05 scaling set).… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-21.05-test.samromur-milljon-test
samromur-milljon (test split only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the official test
split of palli23/samromur-milljon-splits (5017 utt.) -- the training
pool is not redistributed here. Used in the paper as the "Miljon, raw"
zero-shot difficulty anchor (92.5% of its prompts share exact text with… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-milljon-test.samromur_asr
Dataset Card for samromur_asr
Dataset Summary
This is a modfied copy of the dataset from The Language and Voice Laboratory in RU.
This is the first release of the Samrómur Icelandic Speech corpus that contains 100.000 validated utterances.
The corpus is a result of the crowd-sourcing effort run by the Language and Voice Lab at the Reykjavik University, in cooperation with Almannarómur, Center for Language Technology.
Languages
The audio is in Icelandic.
The… See the full description on the dataset page: https://huggingface.co/datasets/DavidErikMollberg/samromur_asr.samromur_syntheticSamrómur Synthetic consists of 72 hours of synthetized speech in Icelandic.
