CoolFace
17 results

Tigrinya

fgaim /tigrinya-abusive-language-detection Tigrinya Abusive Language Detection (TiALD) Dataset TiALD is a large-scale, multi-task benchmark dataset for abusive language detection in the Tigrinya language. It consists of 13,717 YouTube comments annotated for abusiveness, sentiment, and topic tasks. The dataset includes comments written in both the Ge’ez script and prevalent non-standard Latin transliterations to mirror real-world usage. The dataset also includes contextual metadata such as video titles and VLM-generated and… See the full description on the dataset page: https://huggingface.co/datasets/fgaim/tigrinya-abusive-language-detection.text-classification5 likes97 downloads11mo agoHugging Facebadrex /tigrinya-speechaudio10K<n<100K0 likes77 downloads11mo agoHugging FaceHarbidel /tigrinya-asr-mergedgated tigrinya-asr-merged A merged Tigrinya speech-recognition dataset, combining and deduplicating: badrex/tigrinya-speech (train pool) google/WaxalNLP config tir_asr (train pool) UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set) Processing Standardized to audio (16kHz mono) and text columns, with a source column tracking origin Unicode NFC-normalized transcripts, empty transcripts dropped Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.audioautomatic-speech-recognition10K<n<100K0 likes74 downloads25d agoHugging Facefgaim /GLOCR-Tigrinya GLOCR: GeezLab OCR Dataset Overview GLOCR is a Text Recognition (TR) and Optical Character Recognition (OCR) dataset for the Tigrinya language. The dataset contains a total of 661K image-label pairs from multiple data sources. In addition to the characters-only data, the major part of the dataset is a collection of multi-word text images with labels from three categories: News (from Haddas Ertra newspaper), the Bible, and random-trigrams of the 150k most common words in… See the full description on the dataset page: https://huggingface.co/datasets/fgaim/GLOCR-Tigrinya.imageimage-to-text1M<n<10M0 likes53 downloads1y agoHugging FaceProfessor /tigrinya-speech-data Tigrinya Speech Data (Pooled) A ~180.0-hour Tigrinya speech corpus, drawn from a single source (Afrivoice Ethiopia) and filtered to only genuinely transcribed audio. Part of the AfroNet multi-language TTS data effort — sibling release to Yoruba/Hausa/Igbo/Kinyarwanda/Swahili, but Tigrinya (and its four sibling Ethiopian-language releases, Amharic/Oromo/Sidama/Wolaytta) are each published independently, not bundled into one combined "Ethiopia" dataset, even though they share a… See the full description on the dataset page: https://huggingface.co/datasets/Professor/tigrinya-speech-data.text-to-speech10K<n<100K0 likes42 downloads1mo agoHugging Facesaillab /alpaca-tigrinya-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-tigrinya-cleaned.text10K<n<100K0 likes36 downloads2y agoHugging Face