datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FalaBracarense_splitsdataset website: projectofalabracarense
Licence
CC - BY - NC - ND
Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives
cetucconstituicaolapsbmFalAI
Dataset Card for FalAI Dataset
Dataset Summary
The FalAI dataset consists of a total of 265,603 audio files (wav) with associated annotations in the form of metadata.
The FalAI dataset is designed for SLU (Spoken Language Understanding) and is the largest publicly released dataset, in any language, for the task of SLU.
Metadata is available for each recording, including the reference phrase, validation label, user id, demographics such as age, accent, gender, locality and… See the full description on the dataset page: https://huggingface.co/datasets/GTM-UVigo/FalAI.PersianAudiobook
PersianAudiobook
PersianAudiobook is a collection of short Persian audiobook speech clips paired with conservatively refined pseudo-transcriptions. The initial release contains 39,454 accepted examples representing 218.564 hours of mono, 16 kHz speech. It is intended for speech-recognition research, audiobook-domain language modeling, speech representation learning, and carefully reviewed text-to-speech research.
Raw data was collected from IranSeda audiobooks.
The… See the full description on the dataset page: https://huggingface.co/datasets/Pooya-Fallah/PersianAudiobook.coddef
