datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
google-smol-tigre
Google SMOL - Tigre Parallel Corpus Pointer
This repository provides a targeted pointer to the three English-Tigre (en_tig.jsonl) subsets from Google's google/smol dataset across its gatitos, smoldoc, and smolsent tasks.
Dataset Structure
The dataset contains three separate splits mapping directly to the remote source files:
gatitos: Word/phrase-level translation pairs (gatitos/en_tig.jsonl)
smoldoc: Document-level parallel text (smoldoc/en_tig.jsonl)
smolsent:… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/google-smol-tigre.FlowSteer-Dataset
FlowSteer Dataset
A comprehensive evaluation and training benchmark containing 12 evaluation datasets and 1 training dataset across 3 domains: Math, Code, and QA.
Dataset Structure
├── train/ # Training data
│ └── train_12k.jsonl # 12,000 balanced training samples
└── eval/ # Evaluation data
├── gsm8k.jsonl # 128 samples
├── math.jsonl # 128 samples
├── aime2025.jsonl # 30 samples
├──… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/FlowSteer-Dataset.tigre-data-monolingual-text
Tigre Corpus — Two Files
This release is split into two separate files, because they contain two structurally different kinds of text segmentation. Use them accordingly — do not assume every line across both files represents the same kind of unit.
At a glance:
tig_corpus_newspaper_sentences.txt — genuine clause/sentence-level segments, derived from real punctuation in the source newspaper text.
tig_corpus_narrative_chunks.txt — fixed-length 25-token chunks from an embedded… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-monolingual-text.tigre-data-parallel-multilingual
Tigre Parallel Multilingual Dataset (Tigre-Data 1.0)
Overview
This repository introduces the Parallel Multilingual Text component of the Tigre language resource collection. Tigre is an under-resourced South Semitic language within the Afro-Asiatic family.
The goal of Tigre-Data 1.0 is to accelerate research in low-resource NLP and morphologically rich language modeling. This dataset provides a clean, high-quality parallel corpus essential for developing and… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-parallel-multilingual.SkillFlow-Dataset
SkillFlow Dataset
This repository stores the IID training and validation data used by the SkillFlow training code.
Code
The training code is available at:
https://github.com/beita6969/SkillFlow
Files
File
Split
Samples
train_v3.json
train
3500
test_iid_v3.json
iid validation
798
Paper alignment
This release is aligned with the in-distribution benchmark families described in the SkillFlow appendix: HotpotQA, TriviaQA… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/SkillFlow-Dataset.tigre-data-kenLM
Tigre 5-gram Language Model (KenLM)
Overview
This repository provides a 5-gram Language Model (LM) for the Tigre language, trained using the KenLM toolkit. This model is a foundational resource for various downstream NLP and speech applications, including:
Rescoring hypotheses in Automatic Speech Recognition (ASR).
Improving text generation and fluency in Machine Translation (MT).
Performing basic text filtering and quality control.
The model is provided in the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-kenLM.tigre-data-lexicon
Tigre Data Lexicon (tigre-data-lexicon)
Overview
This repository contains the Tigre Data Lexicon, a specialized linguistic resource designed to support the development of speech and language technologies for Tigre, an under-resourced Semitic language. This lexicon serves as a foundational component for bridging the gap between written text and spoken language, facilitating advancements in Artificial Intelligence (AI) and Natural Language Processing (NLP) for the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-lexicon.be_iter02tigre-tts-trainingtigre-speech-text-aligned
Tigre Speech Corpus
1. Overview
This Tigre Speech Corpus is a curated collection of 18,470 aligned audio–text pairs designed to support research and development in speech technologies for Tigre (tig), an under-resourced South Semitic language spoken primarily in Eritrea. The dataset contains approximately 32 hours of recorded speech contributed by over 100 native speakers.
It reflects a collective effort by Tigre-speaking contributors worldwide, including a… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-speech-text-aligned.tigre-hubert-speech
Tigre HuBERT Speech Resources
Self-supervised speech resources for Tigre (ISO 639-3: tig), a Semitic
language spoken primarily in Eritrea and Sudan with very limited existing
speech-technology support. This repository bundles a Tigre-pretrained HuBERT
encoder, a discrete unit-discovery model, forced-aligned transcripts with
word-level unit sequences, and a word-to-unit pseudo-lexicon -- everything
needed to reproduce or extend this work.
Dataset Summary
6777… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-hubert-speech.tigre-data-fasttext
Tigre Word Embedding Models (FastText)
Model Name
Language
Task
License
tig.bin
Tigre (tig)
Word Embeddings (FastText)
CC-BY-SA-4.0
tigre.vec
Tigre (tig)
Word Embeddings (Word2Vec format)
CC-BY-SA-4.0
Overview
This repository introduces the first comprehensive public collection of resources for the Tigre language — an under-resourced South Semitic language within the Afro-Asiatic family. The release aggregates multiple modalities (text + speech)… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-fasttext.tigre-data-wikipedia
Tigre Wikipedia Corpus (tigwiki)
Overview
This repository houses the Tigre Wikipedia Corpus, a foundational linguistic resource containing all non-template articles from https://tig.wikipedia.org.
Tigre is an under-resourced South Semitic language within the Afro-Asiatic family.
This dataset serves as a critical component for bridging the digital divide, facilitating the development of Natural Language Processing (NLP) models—including Language Models (LMs)… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-wikipedia.beit3-keyframe24tigre-tts-modeltigre-hubert-candidate-phones
Draft Candidate Phone Inventory for Tigre — Reviewer Guide
What this is
An automatically-derived candidate sound-unit inventory for Tigre, built
from a self-supervised HuBERT model trained on Tigre audio (Common Voice),
with sounds grouped by unsupervised clustering (k-means) rather than
linguistic analysis.
Two granularities are provided:
candidate_phones_fine.csv — 100 fine-grained clusters. These
likely include allophones (positional/contextual variants of the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-hubert-candidate-phones.beit3_vqa_answer2label.txttigre-data-speech-audio
🇪🇷 Tigre Speech Corpus (Broadcast Audio)
A large-scale, open-source speech dataset for the Tigre language (ISO 639-3: tig), developed to support Automatic Speech Recognition (ASR), speech technology research, and language documentation for one of the least-resourced languages in the Afro-Asiatic family.
This corpus provides hundreds of hours of real-world spoken Tigre, sourced from long-form public radio programming, making it one of the most substantial publicly available… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-speech-audio.beit-prophetnet-processedbeit3_batch2beit-base-patch16-224-in21kbe_iter01
