datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EUbookshop-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the EUbookshop dataset, consisting of 33,634 text segments.
The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 159 hours and 45 minutes (159:45:05) spread across 67,268 utterances.
Dataset Structure
Dataset({
features: ['audio'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/EUbookshop-Speech-Irish.Wikimedia-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the Wikimedia dataset, consisting of 7,545 text segments.
The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 34 hours and 23 minutes (34:23:12) spread across 15,090 utterances.
Dataset Structure
Dataset({
features: ['audio', 'text_ga'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Wikimedia-Speech-Irish.Irodori-Ja-Spk4-10k
SynDataLab/Irodori-Ja-Spk4-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk4): 30s female, news-anchor mature — 30代女性、ニュースキャスター風の落ち着いた声.
How this speaker was made
The voice identity for Spk4 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk4-10k.Irodori-Ja-Spk3-10k
SynDataLab/Irodori-Ja-Spk3-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声.
How this speaker was made
The voice identity for Spk3 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k.Living-Audio-Irish
Dataset Details
Living Audio Irish speech corpus. This version is based on the Irish dataset on Kaggle.
The original dataset with audio in more languages is available on GitHub as part of the Idlak project.
The details of the Irish portion of the Living Audio dataset are as follows:
Speaker
Language
Accent
Gender
Total duration(mm:ss)
Sample rate (Hz)
CLL
Irish (ga)
Non-native (ie)
Man
61:56
48,000
Dataset Structure
Dataset({
features: ['sentence'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Living-Audio-Irish.Tatoeba-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the Tatoeba dataset, consisting of 1,983 text segments.
The dataset consists of two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 2 hours and 39 minutes (02:39:31) spread across 3,966 utterances.
Dataset Structure
Dataset({
features: ['audio', 'text_ga'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Tatoeba-Speech-Irish.Irodori-Ja-Spk1-10k
SynDataLab/Irodori-Ja-Spk1-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk1): 30s male, calm conversational — 30代男性、落ち着いた自然な会話調.
How this speaker was made
The voice identity for Spk1 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk1-10k.Children_Counsel
아동·청소년 상담 데이터셋 (Children Counseling Dataset)
This dataset contains counseling data for children and adolescents, including both audio recordings and transcriptions.
Dataset Structure
The dataset is organized as follows:
audio/: Contains the audio recordings of counseling sessions in MP3 format
data/: Contains JSON files with transcriptions and metadata for each session
Usage
This dataset can be used for:
Training speech recognition models for counseling… See the full description on the dataset page: https://huggingface.co/datasets/ironDong/Children_Counsel.SqCLIRIL
🗣️ SqCLIRIL: Spoken Query Benchmark for Cross-Lingual IR in Indian Languages
SqCLIRIL is a Spoken Query Benchmark designed to evaluate cross-lingual information retrieval (CLIR) systems using both spoken and text queries.It covers five Indian languages — Hindi, Gujarati, Bengali, Kannada, and English — with diverse speech samples from male and female speakers to capture natural variability in pronunciation and acoustic conditions.
📘 Dataset Summary
Feature… See the full description on the dataset page: https://huggingface.co/datasets/irlab-daiict/SqCLIRIL.Irodori-Ja-Spk2-10k
SynDataLab/Irodori-Ja-Spk2-10k
10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos.
Speaker (Spk2): 30s female, narrator-style natural — 30代女性、ナレーター風の自然な声.
How this speaker was made
The voice identity for Spk2 was created in two stages:
Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk2-10k.Irish-Speech-Dataset
Irish Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Irish (ga)
🏷️ Tags
Audio, ML, Machine, Machine Learning, Speech, Speech Recognition, Irish
📦 Size Category
n < 1K
jvsu_asr
Javanese and Sundanese ASR Subset
This dataset provides curated subsets of the OpenSLR SLR35 (Javanese) and SLR36 (Sundanese) speech corpora. Each sample contains a reference transcript along with speaker ID and original archive source.
Dataset Structure
Each language has its own configuration:
javanese
sundanese
For each config, we provide:
train
test
Features
filename — the audio file name (without extension)
userid — speaker/user identifier
label —… See the full description on the dataset page: https://huggingface.co/datasets/irasalsabila/jvsu_asr.irula
Irula Translation Transcription
About
Irula_Translation_Transcription
Tools
We employed Karya and Atekho for collecting and recording data. The complete dataset is transcribed and exported using MATra Lab. Both Atekho and MAtra Lab are part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.
Speakers
Benakkesh, Unais, Aliza, Nithin, Gowri, Roy
Annotators
WorkShop003, WorkShop005, WorkShop004, Work Shop Unreal… See the full description on the dataset page: https://huggingface.co/datasets/project-boli/irula.
