datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libretranslate-en-kab-suggestions
Kabyle Suggestions Dataset
This dataset contains English-to-Kabyle translation suggestions submmitted by users using LibreTranslate, designed to support the development and evaluation of machine translation tools for the Kabyle language.
taqpol_insilico_dms
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: Yulia E. Tomilova,
Nikolai E. Russkikh,
Igor M. Yi,
Elizaveta V. Shaburova,
Viktor N. Tomilov,
Galina B. Pyrinova,
Svetlana O. Brezhneva,
Olga S. Tikhonyuk,
Nadezhda S. Gololobova,
Dmitriy V. Popichenko,
Maxim O. Arkhipov,
Leonid O. Bryzgalov,
Evgeny V. Brenner… See the full description on the dataset page: https://huggingface.co/datasets/nerusskikh/taqpol_insilico_dms.common-voice-scripted-speech-kab-26-huge
Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned)
Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.kabyle-verbs
Kabyle Verbs — Kabyle Verb Conjugation
Kabyle verb conjugation dataset — 6,198 verbs, ~344,000 conjugated forms, covering aorist, preterite, imperative, participles, and intensive forms.
Data source: amyag.com, work by Kamal Nait Zerrad.
Summary
Property
Value
Language
Kabyle (taqbaylit)
Verbs
6,198
Total conjugated forms
344,745
Unique forms
214,276
Grammatical tenses
11 (aorist, preterite, negative preterite, imperative, intensive aorist… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-verbs.cm.trial
Dataset Card for Common Voice Corpus 11.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 24210 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 16413 validated hours in 100 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/taqwa92/cm.trial.Kabyle_Road_Traffic_Code
Kabyle-English Road Traffic Code Dataset
A bilingual parallel corpus of 102 road traffic signs and regulations in English and Kabyle (Taqbaylit), an Amazigh language spoken in Algeria.
Categories
Dangers (Imihiten): Warning signs (39 entries)
Prohibitions (Tigedlin): Prohibitory signs (35 entries)
Obligations (Timariwin): Mandatory signs (16 entries)
End of Restrictions: End of regulation signs (12 entries)
Splits
Split
Size
Train
62… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/Kabyle_Road_Traffic_Code.f5tts-kabyle-dataset
F5-TTS Kabyle Dataset
Clean, deduplicated audio-text dataset for Kabyle (Taqbaylit / Tamaziɣt) TTS fine-tuning with F5-TTS.
Statistics
Metric
Value
Total clips
59,462
Total duration
41.30 hours
Sample rate
24 kHz mono
Avg clip length
2.50s
Min clip length
1.00s
Max clip length
12.65s
Unique phrases
59,462 (0% duplicates)
Unique characters
112
Sources
Tatoeba (67.8%) + Common Voice 26 tiny (32.2%)
Source Datasets… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/f5tts-kabyle-dataset.tatoeba-kabyle-mono-cleaned
tatoeba-kabyle-mono-cleaned
Cleaned and quality-assessed monolingual Kabyle corpus extracted from Tatoeba.
Summary
This dataset contains sentences from Tatoeba tagged as Kabyle (lang == "kab"), processed through a multi-layer linguistic filtering pipeline combining orthographic normalization, language identification (GlotLID v3 + DistilBERT Kabyle/Tachelhit classifier), code-switching detection (MaskLID), and lexical validation (Kabyle Hunspell dictionary).… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-mono-cleaned.mg.trial4mg2_datacm.mgb2kabyle-toponyms
Algeria French–Kabyle Toponym Corpus
A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab.
Description
This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI).
The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.tatoeba-en-kab
Tatoeba English-Kabyle Parallel Corpus
A cleaned and aligned English-Kabyle parallel corpus extracted from Tatoeba, with both direct en↔kab links and indirect kab→fra→en links.
Statistics
Split
Pairs
train
240,056
dev
2,449
test
2,449
Total
244,954
Source
Tatoeba direct en↔kab links
Indirect kab→fra→en links (Kabyle linked to French, French linked to English)
Cleaning Pipeline
Character standardization: Fixed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-en-kab.ayamun-pdfs230 pdf files from Ayamun.
tatoeba-kabyle-audio
Tatoeba Kabyle Audio Dataset
A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected.
Dataset Description
This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.arkavidia-final-fekabyle-corpuskabyle-english-translatewiki
English-Kabyle Parallel Corpus
A clean, deduplicated parallel corpus of English → Kabyle (Taqbaylit) translations extracted from the translatewiki.net bulk dump (2026-01-01).
Dataset Summary
Attribute
Value
Language pair
English (en) → Kabyle (kab)
Total unique pairs
8,871
Source
translatewiki.net
License
CC BY 3.0
Domain
Software localization, UI strings, documentation
Dataset Structure
{
"translation": {
"en": "Hello"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-translatewiki.kab-en-toponyms-sentences
English-Kabyle Parallel Corpus for Machine Translation
This dataset contains 32,024 grammatically flawless parallel sentence pairs mapping English to literary Kabyle (Taqbaylit kab).
This corpus was synthesized using a linguistically-informed morphosyntactic rule engine paired with clean OpenStreetMap toponym registries from boffire/kabyle-toponyms. It handles complex phonetic mutations natively, making it a state-of-the-art bootstrapping asset for fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kab-en-toponyms-sentences.kabyle-named-entities
Kabyle Standardized Named Entities Dataset
This is a manually curated parallel corpus in Kabyle complete with semantic English contextual translations and structured Named Entity Recognition (NER) tag assignments.
Dataset Structure
kabyle_standardized: Target entity string conforming to standardized orthographic regulations.
english_translation: High-context semantic meaning, institutional purpose, or micro-topographic geographical breakdowns.
entity_category:… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-named-entities.EOS-Continued-Pretraining-Dataset
EOS Continued Pre-Training Dataset (Indonesia)
Deskripsi Dataset
EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia yang dikurasi untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM).
Tujuan utama dari dataset ini adalah untuk melakukan Domain Adaptation, yaitu meningkatkan kemampuan model dalam memahami konteks, terminologi, dan nuansa pada dua domain strategis di Indonesia:
Pengawasan Ruang Digital (PRD)
Digital Talent Pool… See the full description on the dataset page: https://huggingface.co/datasets/taqiyudinadn/EOS-Continued-Pretraining-Dataset.nllb_en_kab
NLLB English-Kabyle Parallel Corpus (Filtered & Cleaned)
Parallel English–Kabyle sentence pairs derived from the OPUS-NLLB corpus, filtered with GlotLid v3 and cleaned through a multi-stage Kabyle-specific pipeline.
Dataset Structure
nllb_en_kab.parquet: Parquet file with two columns:
english: English sentence
kabyle: Kabyle sentence
Statistics
Metric
Count
Total sentence pairs
2,786,012
Non-null English
2,786,012
Non-null Kabyle… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/nllb_en_kab.timucuha-kabyle-tales
Timucuha Trilingual Corpus
A parallel corpus of Kabyle (Tamazight) folk tales with French and English translations.
Source
The original Kabyle tales were collected and digitized by the Association Culturelle Numidya.
This dataset is derived from their Timucuha project, which preserves and promotes Kabyle oral tradition.
Website: https://timucuha.numidya.net/
Organization: Association Culturelle Numidya
Dataset Description
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/timucuha-kabyle-tales.rradyu-tis-snat
Rradyu Tis Snat — Kabyle Podcasts from Radio Algérie Chaîne 2
Status: work in progress. This README is a first draft with placeholders
(marked TODO) to fill in as the dataset grows. Metadata above (license,
size_categories) will need updating as the collection is built out.
Dataset Description
This dataset is a collection of Kabyle-language ("Taqbaylit") audio podcasts
from Radio Algérie Chaîne 2
(podcast.radioalgerie.dz),
the Algerian public radio channel… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/rradyu-tis-snat.TA_Quote_Codehumour-taquin-francaiskabyle-english-TM
Kabyle–English Translation Memory
A bilingual translation memory containing 121,725 sentence pairs
in Kabyle (kab) and English (en), built from open-source software
localisation data aggregated through an automated pipeline.
Dataset structure
Each record contains the following fields:
Field
Type
Description
source
string
Source segment (English)
source_lang
string
Always "en"
target
string
Target segment (Kabyle)
target_lang
string
Always "kab"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-TM.kabyle-g2p-training-data
Kabyle G2P Training Data
Phonetically-annotated Kabyle (Taqbaylit) text corpus for training Grapheme-to-Phoneme (G2P) models. Generated using the orthography2ipa rule-based phonemizer for Kabyle.
Dataset Overview
Property
Value
Language
Kabyle (kab) — Afro-Asiatic, Berber
Total pairs
59,462
Source
boffire/kabyle-piper-22khz
Phonemizer
orthography2ipa (dev branch)
IPA standard
Narrow transcription with Kabyle-specific allophony
License
CC0… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-g2p-training-data.cm.trial1qwen-1.5b-blind-spotsTechnical Analysis: Blind Spots of Qwen2.5-1.5B
How the Model was Loaded:
The model was loaded in a Google Colab environment using the transformers library with torch_dtype=torch.bfloat16 to fit within the T4 GPU memory limits.
Fine-tuning Strategy:
To fix the identified logical, grammatical, and instruction-following errors, the model should undergo Supervised Fine-Tuning (SFT). This would involve training the model on "Chain-of-Thought" (CoT) datasets where the model is taught to explain its… See the full description on the dataset page: https://huggingface.co/datasets/Taqiiiiiiiii/qwen-1.5b-blind-spots.
