datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KabTifinagh
KabTifinagh
A standardized bidirectional script transliteration and schwa (e) vowel restoration benchmark for Kabyle (Taqbaylit, kab, Latin & Tifinagh scripts), created by the AƔBALU project.
KabTifinagh normalises, repairs, deduplicates, and structures 497,944 parallel sentence entries matching Neo-Tifinagh (kab_Tfng) to canonical Kabyle Latin (kab_Latn), alongside 123,852 English and 205,637 French trilingual sentence alignments.
from datasets import load_dataset
script =… See the full description on the dataset page: https://huggingface.co/datasets/agbalu/KabTifinagh.KabInflect
KabInflect
A morphological inflection and analysis benchmark for Kabyle (Taqbaylit, kab, Latin script), from the
AƔBALU project.
336,151 inflected verb form entries across 13,226 unique verb lemmas, partitioned into
paradigmatically sealed splits (0 paradigm leakage), plus 6,198 complete verb conjugation tables.
from datasets import load_dataset
inflect = load_dataset("agbalu/KabInflect", "inflection")
analysis = load_dataset("agbalu/KabInflect", "analysis")
paradigms =… See the full description on the dataset page: https://huggingface.co/datasets/agbalu/KabInflect.manas-dataset-v2
Manas Dataset
Statistics
Total clean conversations: 1061
Train: 954
Eval: 107
Format
{
"conversations": [
{"from": "system", "value": "..."},
{"from": "human", "value": "..."},
{"from": "gpt", "value": "..."}
]
}
Kabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.KabLiterary
KabLiterary
A 28,907-paragraph, 657,713-word corpus of classical world literature in canonical Kabyle Latin
orthography, from the AƔBALU project.
Eleven works — translated or originally composed in Kabyle — spanning epic poetry, gothic fiction,
philosophical prose, folktales, military strategy, and 19th-century novels. The corpus is designed
for language modelling, vocabulary probing, fine-tuning, and literary translation benchmarking
where long-form, high-register Kabyle text… See the full description on the dataset page: https://huggingface.co/datasets/agbalu/KabLiterary.wesnoth-ethea-canon-campaignspashto-kabul-treaty-1921-sft
Dataset Card for Pashto Kabul Treaty 1921 SFT
Dataset Summary
This dataset contains the complete Pashto translation of the 1921 Treaty between the British and Afghan Governments (also known as the Kabul Treaty), along with 100 question-answer pairs derived from the treaty text. The original treaty was signed at Kabul on November 22, 1921, and ratifications were exchanged on February 6, 1922.
The dataset is designed for Supervised Fine-Tuning (SFT) of Large… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-kabul-treaty-1921-sft.manas-dataset
Manas AI Training Dataset
This repository contains the official high-quality fine-tuning training dataset for Manas — a warm, culturally-grounded mental health companion AI designed for Indian cultural contexts.
Dataset Overview
The dataset is organized into 11 distinct categories representing psychological pressures and emotional challenges common in urban and semi-urban Indian demographics. Each dialogue is generated using a distinct thematic persona archetype… See the full description on the dataset page: https://huggingface.co/datasets/kabir4756/manas-dataset.
