datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Quran-kabyle-ayt-mensour
Dataset Card: Quran Kabyle Translation (Ramdane At Mensour)
Dataset Summary
This dataset contains the Kabyle (Taqbaylit / Amazigh) translation of the Holy Quran titled "LEQWṚAN S TMAZIƔT", translated by Ramdane At Mensour (Remḍan At Menṣuṛ). It provides verse-by-verse alignments across three script representations: legacy custom-encoded ASCII, standardized INALCO Latin, and IRCAM Tifinagh.
Previously, digital distributions of this translation across mobile apps… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Quran-kabyle-ayt-mensour.Kabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.Kabyle-French
French - Kabyle (Tatoeba)
This dataset contains translation pairs for French (fr) and Kabyle (kab). The data was collected from the Tatoeba Project, a free collaborative online database of example sentences.
⚠️ Important Note on Quality
Disclaimer: This dataset has been exported automatically and has not been manually verified. While Tatoeba relies on community contributions, errors or inconsistencies in translation pairs may exist. Use with appropriate caution.
Kabyle-French-pairs
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/sifalklioui/kabyle-french
tamazight-kabyle-text
Kabyle Tamazight Text Corpus
A dataset of over 3 million Tamazight(kabyle) text lines for NLP tasks.
kabyle-named-entities
Kabyle Standardized Named Entities Dataset
This is a manually curated parallel corpus in Kabyle complete with semantic English contextual translations and structured Named Entity Recognition (NER) tag assignments.
Dataset Structure
kabyle_standardized: Target entity string conforming to standardized orthographic regulations.
english_translation: High-context semantic meaning, institutional purpose, or micro-topographic geographical breakdowns.
entity_category:… See the full description on the dataset page: https://huggingface.co/datasets/boffire/kabyle-named-entities.kabyle-g2p-training-data
Kabyle G2P Training Data
Phonetically-annotated Kabyle (Taqbaylit) text corpus for training Grapheme-to-Phoneme (G2P) models. Generated using the orthography2ipa rule-based phonemizer for Kabyle.
Dataset Overview
Property
Value
Language
Kabyle (kab) — Afro-Asiatic, Berber
Total pairs
59,462
Source
boffire/kabyle-piper-22khz
Phonemizer
orthography2ipa (dev branch)
IPA standard
Narrow transcription with Kabyle-specific allophony
License
CC0… See the full description on the dataset page: https://huggingface.co/datasets/boffire/kabyle-g2p-training-data.
