CoolFace
20 results

yoruba

Africanvoice /African_voices_yorubaaudio100K<n<1M0 likes1k downloads1mo agoHugging FaceMeghanaKap /yoruba_dataset_encodedgatedtext100K<n<1M0 likes963 downloads14d agoHugging Facebytel0rd /yoruba_audio_translatedThis is a copy of odunola/Yoruba_translate_preprocessed, the only difference is, it's already splitted into train & test. Awesome credits to her, her license applies too. audiotranslation10K<n<100K3 likes479 downloads2y agoHugging Facemichsethowusu /yoruba-speech-text-parallel Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba - yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.audioautomatic-speech-recognition1M<n<10M3 likes393 downloads1y agoHugging Faceajesujoba /yoruba_gv_nerThe Yoruba GV NER dataset is a labeled dataset for named entity recognition in Yoruba. The texts were obtained from Yoruba Global Voices News articles https://yo.globalvoices.org/ . We concentrate on four types of named entities: persons [PER], locations [LOC], organizations [ORG], and dates & time [DATE]. The Yoruba GV NER data files contain 2 columns separated by a tab ('\t'). Each word has been put on a separate line and there is an empty line after each sentences i.e the CoNLL format. The first item on each line is a word, the second is the named entity tag. The named entity tags have the format I-TYPE which means that the word is inside a phrase of type TYPE. For every multi-word expression like 'New York', the first word gets a tag B-TYPE and the subsequent words have tags I-TYPE, a word with tag O is not part of a phrase. The dataset is in the BIO tagging scheme. For more details, see https://www.aclweb.org/anthology/2020.lrec-1.335/token-classification1K<n<10K0 likes233 downloads3y agoHugging Faceajesujoba /yoruba_text_c3Yoruba Text C3 is the largest Yoruba texts collected and used to train FastText embeddings in the YorubaTwi Embedding paper: https://www.aclweb.org/anthology/2020.lrec-1.335/text-generation100K<n<1M3 likes228 downloads3y agoHugging Face