datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tigre-data-monolingual-text
Tigre Corpus — Two Files
This release is split into two separate files, because they contain two structurally different kinds of text segmentation. Use them accordingly — do not assume every line across both files represents the same kind of unit.
At a glance:
tig_corpus_newspaper_sentences.txt — genuine clause/sentence-level segments, derived from real punctuation in the source newspaper text.
tig_corpus_narrative_chunks.txt — fixed-length 25-token chunks from an embedded… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-monolingual-text.tigre-data-parallel-multilingual
Tigre Parallel Multilingual Dataset (Tigre-Data 1.0)
Overview
This repository introduces the Parallel Multilingual Text component of the Tigre language resource collection. Tigre is an under-resourced South Semitic language within the Afro-Asiatic family.
The goal of Tigre-Data 1.0 is to accelerate research in low-resource NLP and morphologically rich language modeling. This dataset provides a clean, high-quality parallel corpus essential for developing and… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-parallel-multilingual.SkillFlow-Dataset
SkillFlow Dataset
This repository stores the IID training and validation data used by the SkillFlow training code.
Code
The training code is available at:
https://github.com/beita6969/SkillFlow
Files
File
Split
Samples
train_v3.json
train
3500
test_iid_v3.json
iid validation
798
Paper alignment
This release is aligned with the in-distribution benchmark families described in the SkillFlow appendix: HotpotQA, TriviaQA… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/SkillFlow-Dataset.tigre-data-lexicon
Tigre Data Lexicon (tigre-data-lexicon)
Overview
This repository contains the Tigre Data Lexicon, a specialized linguistic resource designed to support the development of speech and language technologies for Tigre, an under-resourced Semitic language. This lexicon serves as a foundational component for bridging the gap between written text and spoken language, facilitating advancements in Artificial Intelligence (AI) and Natural Language Processing (NLP) for the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-lexicon.tigre-hubert-speech
Tigre HuBERT Speech Resources
Self-supervised speech resources for Tigre (ISO 639-3: tig), a Semitic
language spoken primarily in Eritrea and Sudan with very limited existing
speech-technology support. This repository bundles a Tigre-pretrained HuBERT
encoder, a discrete unit-discovery model, forced-aligned transcripts with
word-level unit sequences, and a word-to-unit pseudo-lexicon -- everything
needed to reproduce or extend this work.
Dataset Summary
6777… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-hubert-speech.beit3_vqa_answer2label.txt
