datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libretranslate-en-kab-suggestions
Kabyle Suggestions Dataset
This dataset contains English-to-Kabyle translation suggestions submmitted by users using LibreTranslate, designed to support the development and evaluation of machine translation tools for the Kabyle language.
TaQAQ-TreeD-1mil
TaQAQ Series
The TaQAQ‑TreeD series is a high‑fidelity, synthetic representation of the NYSE trade and quotes data stream.
Designed as part of the broader TaQAQ series, this collection captures one trillion consecutive trading events—each
comprising detailed trade executions within a fully simulated market environment that adheres to the same temporal granularity,
message formats, and micro‑structure dynamics as real‑world feeds.
DATA : T1mil
ASOF : 251004
MODAL : Nonfused
FORMAT:… See the full description on the dataset page: https://huggingface.co/datasets/bh2821/TaQAQ-TreeD-1mil.plant_classification_v11taqpol_insilico_dms
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: Yulia E. Tomilova,
Nikolai E. Russkikh,
Igor M. Yi,
Elizaveta V. Shaburova,
Viktor N. Tomilov,
Galina B. Pyrinova,
Svetlana O. Brezhneva,
Olga S. Tikhonyuk,
Nadezhda S. Gololobova,
Dmitriy V. Popichenko,
Maxim O. Arkhipov,
Leonid O. Bryzgalov,
Evgeny V. Brenner… See the full description on the dataset page: https://huggingface.co/datasets/nerusskikh/taqpol_insilico_dms.kab-ocr-dataset
OCR Kabyle (Latin + caractères spéciaux)
Dataset synthétique pour entraîner un modèle OCR sur le kabyle.
Structure prévue
images/ : contiendra les images générées
labels.tsv : liste des couples (image → texte)
charset.txt : alphabet utilisé
adlis-pdfscommon-voice-scripted-speech-kab-26-huge
Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned)
Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.kabyle-verbs
Kabyle Verbs — Kabyle Verb Conjugation
Kabyle verb conjugation dataset — 6,198 verbs, ~344,000 conjugated forms, covering aorist, preterite, imperative, participles, and intensive forms.
Data source: amyag.com, work by Kamal Nait Zerrad.
Summary
Property
Value
Language
Kabyle (taqbaylit)
Verbs
6,198
Total conjugated forms
344,745
Unique forms
214,276
Grammatical tenses
11 (aorist, preterite, negative preterite, imperative, intensive aorist… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-verbs.cm.trial
Dataset Card for Common Voice Corpus 11.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 24210 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 16413 validated hours in 100 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/taqwa92/cm.trial.Kabyle_Road_Traffic_Code
Kabyle-English Road Traffic Code Dataset
A bilingual parallel corpus of 102 road traffic signs and regulations in English and Kabyle (Taqbaylit), an Amazigh language spoken in Algeria.
Categories
Dangers (Imihiten): Warning signs (39 entries)
Prohibitions (Tigedlin): Prohibitory signs (35 entries)
Obligations (Timariwin): Mandatory signs (16 entries)
End of Restrictions: End of regulation signs (12 entries)
Splits
Split
Size
Train
62… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/Kabyle_Road_Traffic_Code.f5tts-kabyle-dataset
F5-TTS Kabyle Dataset
Clean, deduplicated audio-text dataset for Kabyle (Taqbaylit / Tamaziɣt) TTS fine-tuning with F5-TTS.
Statistics
Metric
Value
Total clips
59,462
Total duration
41.30 hours
Sample rate
24 kHz mono
Avg clip length
2.50s
Min clip length
1.00s
Max clip length
12.65s
Unique phrases
59,462 (0% duplicates)
Unique characters
112
Sources
Tatoeba (67.8%) + Common Voice 26 tiny (32.2%)
Source Datasets… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/f5tts-kabyle-dataset.bejaia
Béjaïa University Theses Dataset (Taqbaylit / Kabyle)
Description
This dataset contains 647 theses from the Université Abderrahmane Mira de Béjaïa DSpace institutional repository in Algeria:
Folder
Count
Level
master/
640
Master theses (Mémoires de Master)
magister/
7
Magister theses (Mémoires de Magistère)
All documents are in PDF format and were collected from the university's public open-access archive. The theses focus on Kabyle (Taqbaylit)… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/bejaia.weblate-kabyleplant_classification_v2mg.trial4tatoeba-kabyle-mono-cleaned
tatoeba-kabyle-mono-cleaned
Cleaned and quality-assessed monolingual Kabyle corpus extracted from Tatoeba.
Summary
This dataset contains sentences from Tatoeba tagged as Kabyle (lang == "kab"), processed through a multi-layer linguistic filtering pipeline combining orthographic normalization, language identification (GlotLID v3 + DistilBERT Kabyle/Tachelhit classifier), code-switching detection (MaskLID), and lexical validation (Kabyle Hunspell dictionary).… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-mono-cleaned.mg2_datacm.mgb2kabyle-audio-taqbaylit-languagebouira
Bouira University Magistère Theses Dataset
Description
This dataset contains 395 magistère (master's) theses from the Université de Bouira DSpace institutional repository in Algeria.
All documents are in PDF format and were from the university's public open-access archive.
arkavidia-final-fekabyle-toponyms
Algeria French–Kabyle Toponym Corpus
A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab.
Description
This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI).
The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.tatoeba-en-kab
Tatoeba English-Kabyle Parallel Corpus
A cleaned and aligned English-Kabyle parallel corpus extracted from Tatoeba, with both direct en↔kab links and indirect kab→fra→en links.
Statistics
Split
Pairs
train
240,056
dev
2,449
test
2,449
Total
244,954
Source
Tatoeba direct en↔kab links
Indirect kab→fra→en links (Kabyle linked to French, French linked to English)
Cleaning Pipeline
Character standardization: Fixed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-en-kab.ayamun-pdfs230 pdf files from Ayamun.
tatoeba-kabyle-audio
Tatoeba Kabyle Audio Dataset
A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected.
Dataset Description
This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.mg21_datakabyle-corpusEOS-Continued-Pretraining-Dataset
EOS Continued Pre-Training Dataset (Indonesia)
Deskripsi Dataset
EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia yang dikurasi untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM).
Tujuan utama dari dataset ini adalah untuk melakukan Domain Adaptation, yaitu meningkatkan kemampuan model dalam memahami konteks, terminologi, dan nuansa pada dua domain strategis di Indonesia:
Pengawasan Ruang Digital (PRD)
Digital Talent Pool… See the full description on the dataset page: https://huggingface.co/datasets/taqiyudinadn/EOS-Continued-Pretraining-Dataset.mgb2_datakabyle-synth-voice
Kabyle Parallel Corpus (OmniVoice × Tatoeba)
Corpus parallèle de 997 phrases kabyles avec audio généré par OmniVoice.
Statistiques
Langue : kabyle (kab)
Phrases totales : 997
Nouvelles phrases (ce run) : 987
Durée totale : 1958.4s (32.6 min)
Sampling rate : 24000 Hz
Source texte : Tatoeba
Modèle TTS : k2-fsa/OmniVoice
Dernière mise à jour : 2026-05-10T10:49:01.178795
Structure
kabyle_corpus_997/
├── audio/ # Fichiers WAV
├──… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-synth-voice.
