CoolFace
Datasetpublic

taqbaylit/tatoeba-kabyle-audio

Tatoeba Kabyle Audio Dataset A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected. Dataset Description This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes16downloads

taqbaylit/tatoeba-kabyle-audio · main · files are served by the source, never re-hosted here