CoolFace
Datasetpublic

hadamard-2/leyu-amharic-gojjam-dialect

Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.

sourceHugging Facecc-by-4.0updated 14d agoView on Hugging Face
0likes212downloads
4 commits on main
77ae95e14d ago

docs: document the published split.json in the card

hadamard-2
a90c1ff19d ago

data: add leakage-free train/dev/test split assignment

hadamard-2
5bd1b3b20d ago

docs: correct card figures to match the published parquet

hadamard-2
05f66d41mo ago

Duplicate from gheero-Leyu/leyu-amharic-gojjam-dialect

hadamard-2, gheero-Leyu