datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EUbookshop-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the EUbookshop dataset, consisting of 33,634 text segments.
The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 159 hours and 45 minutes (159:45:05) spread across 67,268 utterances.
Dataset Structure
Dataset({
features: ['audio'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/EUbookshop-Speech-Irish.Wikimedia-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the Wikimedia dataset, consisting of 7,545 text segments.
The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 34 hours and 23 minutes (34:23:12) spread across 15,090 utterances.
Dataset Structure
Dataset({
features: ['audio', 'text_ga'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Wikimedia-Speech-Irish.Living-Audio-Irish
Dataset Details
Living Audio Irish speech corpus. This version is based on the Irish dataset on Kaggle.
The original dataset with audio in more languages is available on GitHub as part of the Idlak project.
The details of the Irish portion of the Living Audio dataset are as follows:
Speaker
Language
Accent
Gender
Total duration(mm:ss)
Sample rate (Hz)
CLL
Irish (ga)
Non-native (ie)
Man
61:56
48,000
Dataset Structure
Dataset({
features: ['sentence'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Living-Audio-Irish.Tatoeba-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the Tatoeba dataset, consisting of 1,983 text segments.
The dataset consists of two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 2 hours and 39 minutes (02:39:31) spread across 3,966 utterances.
Dataset Structure
Dataset({
features: ['audio', 'text_ga'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Tatoeba-Speech-Irish.Irish-Speech-Dataset
Irish Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Irish (ga)
🏷️ Tags
Audio, ML, Machine, Machine Learning, Speech, Speech Recognition, Irish
📦 Size Category
n < 1K
