datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libretranslate-en-kab-suggestions
Kabyle Suggestions Dataset
This dataset contains English-to-Kabyle translation suggestions submmitted by users using LibreTranslate, designed to support the development and evaluation of machine translation tools for the Kabyle language.
adlis-pdfscommon-voice-scripted-speech-kab-26-huge
Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned)
Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.kabyle-verbs
Kabyle Verbs — Kabyle Verb Conjugation
Kabyle verb conjugation dataset — 6,198 verbs, ~344,000 conjugated forms, covering aorist, preterite, imperative, participles, and intensive forms.
Data source: amyag.com, work by Kamal Nait Zerrad.
Summary
Property
Value
Language
Kabyle (taqbaylit)
Verbs
6,198
Total conjugated forms
344,745
Unique forms
214,276
Grammatical tenses
11 (aorist, preterite, negative preterite, imperative, intensive aorist… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-verbs.Kabyle_Road_Traffic_Code
Kabyle-English Road Traffic Code Dataset
A bilingual parallel corpus of 102 road traffic signs and regulations in English and Kabyle (Taqbaylit), an Amazigh language spoken in Algeria.
Categories
Dangers (Imihiten): Warning signs (39 entries)
Prohibitions (Tigedlin): Prohibitory signs (35 entries)
Obligations (Timariwin): Mandatory signs (16 entries)
End of Restrictions: End of regulation signs (12 entries)
Splits
Split
Size
Train
62… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/Kabyle_Road_Traffic_Code.f5tts-kabyle-dataset
F5-TTS Kabyle Dataset
Clean, deduplicated audio-text dataset for Kabyle (Taqbaylit / Tamaziɣt) TTS fine-tuning with F5-TTS.
Statistics
Metric
Value
Total clips
59,462
Total duration
41.30 hours
Sample rate
24 kHz mono
Avg clip length
2.50s
Min clip length
1.00s
Max clip length
12.65s
Unique phrases
59,462 (0% duplicates)
Unique characters
112
Sources
Tatoeba (67.8%) + Common Voice 26 tiny (32.2%)
Source Datasets… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/f5tts-kabyle-dataset.bejaia
Béjaïa University Theses Dataset (Taqbaylit / Kabyle)
Description
This dataset contains 647 theses from the Université Abderrahmane Mira de Béjaïa DSpace institutional repository in Algeria:
Folder
Count
Level
master/
640
Master theses (Mémoires de Master)
magister/
7
Magister theses (Mémoires de Magistère)
All documents are in PDF format and were collected from the university's public open-access archive. The theses focus on Kabyle (Taqbaylit)… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/bejaia.weblate-kabyletatoeba-kabyle-mono-cleaned
tatoeba-kabyle-mono-cleaned
Cleaned and quality-assessed monolingual Kabyle corpus extracted from Tatoeba.
Summary
This dataset contains sentences from Tatoeba tagged as Kabyle (lang == "kab"), processed through a multi-layer linguistic filtering pipeline combining orthographic normalization, language identification (GlotLID v3 + DistilBERT Kabyle/Tachelhit classifier), code-switching detection (MaskLID), and lexical validation (Kabyle Hunspell dictionary).… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-mono-cleaned.kabyle-audio-taqbaylit-languagebouira
Bouira University Magistère Theses Dataset
Description
This dataset contains 395 magistère (master's) theses from the Université de Bouira DSpace institutional repository in Algeria.
All documents are in PDF format and were from the university's public open-access archive.
kabyle-toponyms
Algeria French–Kabyle Toponym Corpus
A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab.
Description
This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI).
The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.ayamun-pdfs230 pdf files from Ayamun.
tatoeba-en-kab
Tatoeba English-Kabyle Parallel Corpus
A cleaned and aligned English-Kabyle parallel corpus extracted from Tatoeba, with both direct en↔kab links and indirect kab→fra→en links.
Statistics
Split
Pairs
train
240,056
dev
2,449
test
2,449
Total
244,954
Source
Tatoeba direct en↔kab links
Indirect kab→fra→en links (Kabyle linked to French, French linked to English)
Cleaning Pipeline
Character standardization: Fixed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-en-kab.kabyle-synth-voice
Kabyle Parallel Corpus (OmniVoice × Tatoeba)
Corpus parallèle de 997 phrases kabyles avec audio généré par OmniVoice.
Statistiques
Langue : kabyle (kab)
Phrases totales : 997
Nouvelles phrases (ce run) : 987
Durée totale : 1958.4s (32.6 min)
Sampling rate : 24000 Hz
Source texte : Tatoeba
Modèle TTS : k2-fsa/OmniVoice
Dernière mise à jour : 2026-05-10T10:49:01.178795
Structure
kabyle_corpus_997/
├── audio/ # Fichiers WAV
├──… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-synth-voice.kab-en-toponyms-sentences
English-Kabyle Parallel Corpus for Machine Translation
This dataset contains 32,024 grammatically flawless parallel sentence pairs mapping English to literary Kabyle (Taqbaylit kab).
This corpus was synthesized using a linguistically-informed morphosyntactic rule engine paired with clean OpenStreetMap toponym registries from boffire/kabyle-toponyms. It handles complex phonetic mutations natively, making it a state-of-the-art bootstrapping asset for fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kab-en-toponyms-sentences.tatoeba-kabyle-audio
Tatoeba Kabyle Audio Dataset
A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected.
Dataset Description
This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.kabyle-corpuskabyle-named-entities
Kabyle Standardized Named Entities Dataset
This is a manually curated parallel corpus in Kabyle complete with semantic English contextual translations and structured Named Entity Recognition (NER) tag assignments.
Dataset Structure
kabyle_standardized: Target entity string conforming to standardized orthographic regulations.
english_translation: High-context semantic meaning, institutional purpose, or micro-topographic geographical breakdowns.
entity_category:… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-named-entities.kabyle-english-translatewiki
English-Kabyle Parallel Corpus
A clean, deduplicated parallel corpus of English → Kabyle (Taqbaylit) translations extracted from the translatewiki.net bulk dump (2026-01-01).
Dataset Summary
Attribute
Value
Language pair
English (en) → Kabyle (kab)
Total unique pairs
8,871
Source
translatewiki.net
License
CC BY 3.0
Domain
Software localization, UI strings, documentation
Dataset Structure
{
"translation": {
"en": "Hello"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-translatewiki.timucuha-kabyle-tales
Timucuha Trilingual Corpus
A parallel corpus of Kabyle (Tamazight) folk tales with French and English translations.
Source
The original Kabyle tales were collected and digitized by the Association Culturelle Numidya.
This dataset is derived from their Timucuha project, which preserves and promotes Kabyle oral tradition.
Website: https://timucuha.numidya.net/
Organization: Association Culturelle Numidya
Dataset Description
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/timucuha-kabyle-tales.nllb_en_kab
NLLB English-Kabyle Parallel Corpus (Filtered & Cleaned)
Parallel English–Kabyle sentence pairs derived from the OPUS-NLLB corpus, filtered with GlotLid v3 and cleaned through a multi-stage Kabyle-specific pipeline.
Dataset Structure
nllb_en_kab.parquet: Parquet file with two columns:
english: English sentence
kabyle: Kabyle sentence
Statistics
Metric
Count
Total sentence pairs
2,786,012
Non-null English
2,786,012
Non-null Kabyle… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/nllb_en_kab.kabyle-english-TM
Kabyle–English Translation Memory
A bilingual translation memory containing 121,725 sentence pairs
in Kabyle (kab) and English (en), built from open-source software
localisation data aggregated through an automated pipeline.
Dataset structure
Each record contains the following fields:
Field
Type
Description
source
string
Source segment (English)
source_lang
string
Always "en"
target
string
Target segment (Kabyle)
target_lang
string
Always "kab"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-TM.rradyu-tis-snat
Rradyu Tis Snat — Kabyle Podcasts from Radio Algérie Chaîne 2
Status: work in progress. This README is a first draft with placeholders
(marked TODO) to fill in as the dataset grows. Metadata above (license,
size_categories) will need updating as the collection is built out.
Dataset Description
This dataset is a collection of Kabyle-language ("Taqbaylit") audio podcasts
from Radio Algérie Chaîne 2
(podcast.radioalgerie.dz),
the Algerian public radio channel… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/rradyu-tis-snat.Tughalin_n_Weqcic_Ijahensynthetic-audio-draftskabyle-g2p-training-data
Kabyle G2P Training Data
Phonetically-annotated Kabyle (Taqbaylit) text corpus for training Grapheme-to-Phoneme (G2P) models. Generated using the orthography2ipa rule-based phonemizer for Kabyle.
Dataset Overview
Property
Value
Language
Kabyle (kab) — Afro-Asiatic, Berber
Total pairs
59,462
Source
boffire/kabyle-piper-22khz
Phonemizer
orthography2ipa (dev branch)
IPA standard
Narrow transcription with Kabyle-specific allophony
License
CC0… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-g2p-training-data.kabyle-lyrics
Kabyle lyrics
Here soon.
Meanwhile, listen to Ɛli Ɛemran here : https://www.youtube.com/watch?v=01bK37AP6Y4
Or Silya Uld Muḥend here : https://www.youtube.com/watch?v=klbwQKm1vEM
hunspell-kab
Kabyle Hunspell Dictionary (Cleaned)
A cleaned and documented version of M. Belkacem's Imseɣti n tira n teqbaylit (v1.0, MIT), the only comprehensive open-source spell-checker for the Kabyle language (Taqbaylit, ISO 639-1 kab).
📦 Dataset Contents
File
Description
Size
kab.dic
Cleaned word list with morphological metadata
~685 KB
kab.aff
Original affixation rules (Belkacem v1.0)
~19 KB
cleaning-report.md
Full cleaning audit
~6 KB
tag-vocabulary.md… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/hunspell-kab.
