Tigrinya
Datasets
All datasets matching “Tigrinya”tigrinya-abusive-language-detection
Tigrinya Abusive Language Detection (TiALD) Dataset
TiALD is a large-scale, multi-task benchmark dataset for abusive language detection in the Tigrinya language. It consists of 13,717 YouTube comments annotated for abusiveness, sentiment, and topic tasks. The dataset includes comments written in both the Ge’ez script and prevalent non-standard Latin transliterations to mirror real-world usage.
The dataset also includes contextual metadata such as video titles and VLM-generated and… See the full description on the dataset page: https://huggingface.co/datasets/fgaim/tigrinya-abusive-language-detection.tigrinya-speechtigrinya-asr-merged
tigrinya-asr-merged
A merged Tigrinya speech-recognition dataset, combining and deduplicating:
badrex/tigrinya-speech (train pool)
google/WaxalNLP config tir_asr (train pool)
UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set)
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.GLOCR-Tigrinya
GLOCR: GeezLab OCR Dataset
Overview
GLOCR is a Text Recognition (TR) and Optical Character Recognition (OCR) dataset for the Tigrinya language. The dataset contains a total of 661K image-label pairs from multiple data sources. In addition to the characters-only data, the major part of the dataset is a collection of multi-word text images with labels from three categories: News (from Haddas Ertra newspaper), the Bible, and random-trigrams of the 150k most common words in… See the full description on the dataset page: https://huggingface.co/datasets/fgaim/GLOCR-Tigrinya.tigrinya-speech-data
Tigrinya Speech Data (Pooled)
A ~180.0-hour Tigrinya speech corpus, drawn from a single source (Afrivoice
Ethiopia) and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort — sibling release to Yoruba/Hausa/Igbo/Kinyarwanda/Swahili, but Tigrinya (and
its four sibling Ethiopian-language releases, Amharic/Oromo/Sidama/Wolaytta) are
each published independently, not bundled into one combined "Ethiopia" dataset,
even though they share a… See the full description on the dataset page: https://huggingface.co/datasets/Professor/tigrinya-speech-data.alpaca-tigrinya-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-tigrinya-cleaned.
