datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cm.trial
Dataset Card for Common Voice Corpus 11.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 24210 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 16413 validated hours in 100 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/taqwa92/cm.trial.tatoeba-kabyle-audio
Tatoeba Kabyle Audio Dataset
A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected.
Dataset Description
This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.rradyu-tis-snat
Rradyu Tis Snat — Kabyle Podcasts from Radio Algérie Chaîne 2
Status: work in progress. This README is a first draft with placeholders
(marked TODO) to fill in as the dataset grows. Metadata above (license,
size_categories) will need updating as the collection is built out.
Dataset Description
This dataset is a collection of Kabyle-language ("Taqbaylit") audio podcasts
from Radio Algérie Chaîne 2
(podcast.radioalgerie.dz),
the Algerian public radio channel… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/rradyu-tis-snat.
