datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ruv_tv_unknown_speakersDataset copied from http://hdl.handle.net/20.500.12537/191 by Reykjavik University.
Information can be found at that link.
RUV TV unknown speakers
About the RUV TV unknown speakers corpus
The RUV TV unknown speakers corpus is 281 hours of TV data from six RÚV TV
shows. The data continas 221,759 utterrances from various unlabelled speakers.
The text is normalized. The data is aligned and segmented, ready for ASR
training. Audio conditions vary between recordings. This data set is… See the full description on the dataset page: https://huggingface.co/datasets/tiro-is/ruv_tv_unknown_speakers.ismus
Gamli: Icelandic Oral History Corpus (a.k.a. ismus)
Gamli is an ASR corpus for Icelandic oral histories, the first of its kind for this language, derived from the ethnographic collection of the Árni Magnússon Institute for Icelandic Studies (available on ismus.is) and is the result of collaboration between that same institute and the Icelandic language technology company Tiro. The corpus contains 146 hours of transcribed audio broken down into:
Training set:
∼ 102 hours from… See the full description on the dataset page: https://huggingface.co/datasets/tiro-is/ismus.tiringatiranovoiceTirretirsotiringa
