datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kabyle-verbs
Kabyle Verbs — Kabyle Verb Conjugation
Kabyle verb conjugation dataset — 6,198 verbs, ~344,000 conjugated forms, covering aorist, preterite, imperative, participles, and intensive forms.
Data source: amyag.com, work by Kamal Nait Zerrad.
Summary
Property
Value
Language
Kabyle (taqbaylit)
Verbs
6,198
Total conjugated forms
344,745
Unique forms
214,276
Grammatical tenses
11 (aorist, preterite, negative preterite, imperative, intensive aorist… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-verbs.kabyle-piper-22khzkabyle-corpus-ubouira
Kabyle Paragraph Corpus (Bouira University-DSpace)
336,982 Kabyle sentences extracted from open-access PDFs published onBouira University – the institutional repository ofBouira University (Algeria).
Documents originate from theFaculté des Lettres et des Langues / Département de Langue et Culture Amazighes.
Fields
id : unique identifier
text : paragraph text (UTF-8)
folder : origin sub-corpus
line_no : 1-based line number in original file
Splits… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/kabyle-corpus-ubouira.kabyle_asrMultilang-to-Kabyle-Translationkabyle-corpus-ummto
Kabyle Paragraph Corpus (UMMTO-DSpace)
690,917 Kabyle sentences extracted from open-access PDFs published onDSpace UMMTO – the institutional repository ofUniversité Mouloud Mammeri de Tizi-Ouzou (Algeria).
Documents originate from theFaculté des Lettres et des Langues / Département de Langue et Culture Amazighes.
Fields
id : unique identifier
text : paragraph text (UTF-8)
folder : origin sub-corpus
line_no : 1-based line number in original file
Splits… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/kabyle-corpus-ummto.f5tts-kabyle-dataset
F5-TTS Kabyle Dataset
Clean, deduplicated audio-text dataset for Kabyle (Taqbaylit / Tamaziɣt) TTS fine-tuning with F5-TTS.
Statistics
Metric
Value
Total clips
59,462
Total duration
41.30 hours
Sample rate
24 kHz mono
Avg clip length
2.50s
Min clip length
1.00s
Max clip length
12.65s
Unique phrases
59,462 (0% duplicates)
Unique characters
112
Sources
Tatoeba (67.8%) + Common Voice 26 tiny (32.2%)
Source Datasets… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/f5tts-kabyle-dataset.kabyle-corpus-hca
Kabyle Paragraph Corpus (HCA Algeria)
188,620 Kabyle sentences extracted from open-access PDFs published on HCA Algeria – the institutional repository of Haut Commissariat à l'Amazighité (Algeria).
Documents originate from Haut Commissariat à l'Amazighité.
Fields
id : unique identifier
text : paragraph text (UTF-8)
folder : origin sub-corpus
line_no : 1-based line number in original file
Splits
split
# records
train
169,758
validation
9,431… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/kabyle-corpus-hca.Kabyle_ASR-En_Translationtatoeba-kabyle-mono-cleaned
tatoeba-kabyle-mono-cleaned
Cleaned and quality-assessed monolingual Kabyle corpus extracted from Tatoeba.
Summary
This dataset contains sentences from Tatoeba tagged as Kabyle (lang == "kab"), processed through a multi-layer linguistic filtering pipeline combining orthographic normalization, language identification (GlotLID v3 + DistilBERT Kabyle/Tachelhit classifier), code-switching detection (MaskLID), and lexical validation (Kabyle Hunspell dictionary).… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-mono-cleaned.Kabyle_ASR-Fr_Translationkabyle-emotions-corpus
Kabyle Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Kabyle for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total samples: 4… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kabyle-emotions-corpus.english-kabyle_sentence-pairs_mt560
English-Kabyle Parallel Dataset
This dataset contains parallel sentences in English and Kabyle (Algeria).
Dataset Information
Language Pair: English ↔ Kabyle
Language Code: kab
Country: Algeria
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-kabyle_sentence-pairs_mt560.kabyle-sentiments-corpus
Kabyle Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Kabyle for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 4,887
Positive sentiment: 2789 (57.1%)
Negative… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kabyle-sentiments-corpus.french-kabyle_sentence-pairs
French-Kabyle_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: French-Kabyle_Sentence-Pairs
File Size: 107508244 bytes
Languages: French, French
Dataset Description
The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/french-kabyle_sentence-pairs.english-kabyle_sentence-pairs
English-Kabyle_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: English-Kabyle_Sentence-Pairs
File Size: 139179277 bytes
Languages: English, English
Dataset Description
The dataset contains sentence pairs… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-kabyle_sentence-pairs.kabyle-toponyms
Algeria French–Kabyle Toponym Corpus
A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab.
Description
This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI).
The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.tatoeba-kabyle-audio
Tatoeba Kabyle Audio Dataset
A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected.
Dataset Description
This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.dataset_kabyle
Dataset Card for "dataset_kabyle"
More Information needed
kabyle-english-translatewiki
English-Kabyle Parallel Corpus
A clean, deduplicated parallel corpus of English → Kabyle (Taqbaylit) translations extracted from the translatewiki.net bulk dump (2026-01-01).
Dataset Summary
Attribute
Value
Language pair
English (en) → Kabyle (kab)
Total unique pairs
8,871
Source
translatewiki.net
License
CC BY 3.0
Domain
Software localization, UI strings, documentation
Dataset Structure
{
"translation": {
"en": "Hello"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-translatewiki.Kabyle_TTSkabyle-english-TM
Kabyle–English Translation Memory
A bilingual translation memory containing 121,725 sentence pairs
in Kabyle (kab) and English (en), built from open-source software
localisation data aggregated through an automated pipeline.
Dataset structure
Each record contains the following fields:
Field
Type
Description
source
string
Source segment (English)
source_lang
string
Always "en"
target
string
Target segment (Kabyle)
target_lang
string
Always "kab"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-TM.
