CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taqbaylit /libretranslate-en-kab-suggestions Kabyle Suggestions Dataset This dataset contains English-to-Kabyle translation suggestions submmitted by users using LibreTranslate, designed to support the development and evaluation of machine translation tools for the Kabyle language. texttranslationn<1K0 likes1.7k downloads4mo agoHugging Face02taqbaylit /adlis-pdfsdocumentn<1K1 likes113 downloads5mo agoHugging Face03taqbaylit /common-voice-scripted-speech-kab-26-huge Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned) Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips. Source Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12) Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective) License: CC0-1.0 Generated: 2026-07-12 Cleaning Pipeline Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.audio100K<n<1M0 likes84 downloads2mo agoHugging Face04taqbaylit /kabyle-verbs Kabyle Verbs — Kabyle Verb Conjugation Kabyle verb conjugation dataset — 6,198 verbs, ~344,000 conjugated forms, covering aorist, preterite, imperative, participles, and intensive forms. Data source: amyag.com, work by Kamal Nait Zerrad. Summary Property Value Language Kabyle (taqbaylit) Verbs 6,198 Total conjugated forms 344,745 Unique forms 214,276 Grammatical tenses 11 (aorist, preterite, negative preterite, imperative, intensive aorist… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-verbs.text100K<n<1M0 likes83 downloads3mo agoHugging Face05taqbaylit /Kabyle_Road_Traffic_Code Kabyle-English Road Traffic Code Dataset A bilingual parallel corpus of 102 road traffic signs and regulations in English and Kabyle (Taqbaylit), an Amazigh language spoken in Algeria. Categories Dangers (Imihiten): Warning signs (39 entries) Prohibitions (Tigedlin): Prohibitory signs (35 entries) Obligations (Timariwin): Mandatory signs (16 entries) End of Restrictions: End of regulation signs (12 entries) Splits Split Size Train 62… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/Kabyle_Road_Traffic_Code.tabularn<1K0 likes39 downloads5mo agoHugging Face06taqbaylit /f5tts-kabyle-dataset F5-TTS Kabyle Dataset Clean, deduplicated audio-text dataset for Kabyle (Taqbaylit / Tamaziɣt) TTS fine-tuning with F5-TTS. Statistics Metric Value Total clips 59,462 Total duration 41.30 hours Sample rate 24 kHz mono Avg clip length 2.50s Min clip length 1.00s Max clip length 12.65s Unique phrases 59,462 (0% duplicates) Unique characters 112 Sources Tatoeba (67.8%) + Common Voice 26 tiny (32.2%) Source Datasets… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/f5tts-kabyle-dataset.text10K<n<100K0 likes35 downloads2mo agoHugging Face07taqbaylit /bejaia Béjaïa University Theses Dataset (Taqbaylit / Kabyle) Description This dataset contains 647 theses from the Université Abderrahmane Mira de Béjaïa DSpace institutional repository in Algeria: Folder Count Level master/ 640 Master theses (Mémoires de Master) magister/ 7 Magister theses (Mémoires de Magistère) All documents are in PDF format and were collected from the university's public open-access archive. The theses focus on Kabyle (Taqbaylit)… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/bejaia.documentn<1K0 likes29 downloads4mo agoHugging Face08taqbaylit /weblate-kabyle0 likes29 downloads10d agoHugging Face09taqbaylit /tatoeba-kabyle-mono-cleaned tatoeba-kabyle-mono-cleaned Cleaned and quality-assessed monolingual Kabyle corpus extracted from Tatoeba. Summary This dataset contains sentences from Tatoeba tagged as Kabyle (lang == "kab"), processed through a multi-layer linguistic filtering pipeline combining orthographic normalization, language identification (GlotLID v3 + DistilBERT Kabyle/Tachelhit classifier), code-switching detection (MaskLID), and lexical validation (Kabyle Hunspell dictionary).… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-mono-cleaned.tabular100K<n<1M0 likes20 downloads2mo agoHugging Face10azrunguraya /kabyle-audio-taqbaylit-languageaudion<1K0 likes15 downloads1y agoHugging Face11taqbaylit /bouira Bouira University Magistère Theses Dataset Description This dataset contains 395 magistère (master's) theses from the Université de Bouira DSpace institutional repository in Algeria. All documents are in PDF format and were from the university's public open-access archive. documentn<1K1 likes13 downloads5mo agoHugging Face12taqbaylit /kabyle-toponyms Algeria French–Kabyle Toponym Corpus A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab. Description This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI). The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.tabulartranslation1K<n<10K0 likes13 downloads4mo agoHugging Face13taqbaylit /ayamun-pdfs230 pdf files from Ayamun. documentn<1K1 likes12 downloads4mo agoHugging Face14taqbaylit /tatoeba-en-kab Tatoeba English-Kabyle Parallel Corpus A cleaned and aligned English-Kabyle parallel corpus extracted from Tatoeba, with both direct en↔kab links and indirect kab→fra→en links. Statistics Split Pairs train 240,056 dev 2,449 test 2,449 Total 244,954 Source Tatoeba direct en↔kab links Indirect kab→fra→en links (Kabyle linked to French, French linked to English) Cleaning Pipeline Character standardization: Fixed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-en-kab.texttranslation100K<n<1M0 likes10 downloads3mo agoHugging Face15taqbaylit /kabyle-synth-voice Kabyle Parallel Corpus (OmniVoice × Tatoeba) Corpus parallèle de 997 phrases kabyles avec audio généré par OmniVoice. Statistiques Langue : kabyle (kab) Phrases totales : 997 Nouvelles phrases (ce run) : 987 Durée totale : 1958.4s (32.6 min) Sampling rate : 24000 Hz Source texte : Tatoeba Modèle TTS : k2-fsa/OmniVoice Dernière mise à jour : 2026-05-10T10:49:01.178795 Structure kabyle_corpus_997/ ├── audio/ # Fichiers WAV ├──… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-synth-voice.0 likes9 downloads5mo agoHugging Face16taqbaylit /kab-en-toponyms-sentences English-Kabyle Parallel Corpus for Machine Translation This dataset contains 32,024 grammatically flawless parallel sentence pairs mapping English to literary Kabyle (Taqbaylit kab). This corpus was synthesized using a linguistically-informed morphosyntactic rule engine paired with clean OpenStreetMap toponym registries from boffire/kabyle-toponyms. It handles complex phonetic mutations natively, making it a state-of-the-art bootstrapping asset for fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kab-en-toponyms-sentences.text10K<n<100K0 likes9 downloads4mo agoHugging Face17taqbaylit /tatoeba-kabyle-audio Tatoeba Kabyle Audio Dataset A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected. Dataset Description This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.audioautomatic-speech-recognition10K<n<100K0 likes9 downloads3mo agoHugging Face18taqbaylit /kabyle-corpustext100K<n<1M0 likes8 downloads4mo agoHugging Face19taqbaylit /kabyle-named-entities Kabyle Standardized Named Entities Dataset This is a manually curated parallel corpus in Kabyle complete with semantic English contextual translations and structured Named Entity Recognition (NER) tag assignments. Dataset Structure kabyle_standardized: Target entity string conforming to standardized orthographic regulations. english_translation: High-context semantic meaning, institutional purpose, or micro-topographic geographical breakdowns. entity_category:… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-named-entities.texttoken-classificationn<1K0 likes8 downloads4mo agoHugging Face20taqbaylit /kabyle-english-translatewiki English-Kabyle Parallel Corpus A clean, deduplicated parallel corpus of English → Kabyle (Taqbaylit) translations extracted from the translatewiki.net bulk dump (2026-01-01). Dataset Summary Attribute Value Language pair English (en) → Kabyle (kab) Total unique pairs 8,871 Source translatewiki.net License CC BY 3.0 Domain Software localization, UI strings, documentation Dataset Structure { "translation": { "en": "Hello"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-translatewiki.texttranslation1K<n<10K0 likes5 downloads4mo agoHugging Face21taqbaylit /timucuha-kabyle-tales Timucuha Trilingual Corpus A parallel corpus of Kabyle (Tamazight) folk tales with French and English translations. Source The original Kabyle tales were collected and digitized by the Association Culturelle Numidya. This dataset is derived from their Timucuha project, which preserves and promotes Kabyle oral tradition. Website: https://timucuha.numidya.net/ Organization: Association Culturelle Numidya Dataset Description Property Value… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/timucuha-kabyle-tales.texttranslationn<1K0 likes5 downloads4mo agoHugging Face22taqbaylit /nllb_en_kab NLLB English-Kabyle Parallel Corpus (Filtered & Cleaned) Parallel English–Kabyle sentence pairs derived from the OPUS-NLLB corpus, filtered with GlotLid v3 and cleaned through a multi-stage Kabyle-specific pipeline. Dataset Structure nllb_en_kab.parquet: Parquet file with two columns: english: English sentence kabyle: Kabyle sentence Statistics Metric Count Total sentence pairs 2,786,012 Non-null English 2,786,012 Non-null Kabyle… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/nllb_en_kab.texttranslation1M<n<10M0 likes5 downloads3mo agoHugging Face23taqbaylit /kabyle-english-TM Kabyle–English Translation Memory A bilingual translation memory containing 121,725 sentence pairs in Kabyle (kab) and English (en), built from open-source software localisation data aggregated through an automated pipeline. Dataset structure Each record contains the following fields: Field Type Description source string Source segment (English) source_lang string Always "en" target string Target segment (Kabyle) target_lang string Always "kab"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-TM.texttranslation100K<n<1M0 likes4 downloads2mo agoHugging Face24taqbaylit /rradyu-tis-snat Rradyu Tis Snat — Kabyle Podcasts from Radio Algérie Chaîne 2 Status: work in progress. This README is a first draft with placeholders (marked TODO) to fill in as the dataset grows. Metadata above (license, size_categories) will need updating as the collection is built out. Dataset Description This dataset is a collection of Kabyle-language ("Taqbaylit") audio podcasts from Radio Algérie Chaîne 2 (podcast.radioalgerie.dz), the Algerian public radio channel… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/rradyu-tis-snat.audioautomatic-speech-recognitionn<1K0 likes4 downloads2mo agoHugging Face25taqbaylit /Tughalin_n_Weqcic_Ijahenaudion<1K0 likes3 downloads5mo agoHugging Face26taqbaylit /synthetic-audio-draftsaudion<1K0 likes3 downloads3mo agoHugging Face27taqbaylit /kabyle-g2p-training-data Kabyle G2P Training Data Phonetically-annotated Kabyle (Taqbaylit) text corpus for training Grapheme-to-Phoneme (G2P) models. Generated using the orthography2ipa rule-based phonemizer for Kabyle. Dataset Overview Property Value Language Kabyle (kab) — Afro-Asiatic, Berber Total pairs 59,462 Source boffire/kabyle-piper-22khz Phonemizer orthography2ipa (dev branch) IPA standard Narrow transcription with Kabyle-specific allophony License CC0… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-g2p-training-data.texttext-to-speech10K<n<100K1 likes3 downloads2mo agoHugging Face28taqbaylit /kabyle-lyrics Kabyle lyrics Here soon. Meanwhile, listen to Ɛli Ɛemran here : https://www.youtube.com/watch?v=01bK37AP6Y4 Or Silya Uld Muḥend here : https://www.youtube.com/watch?v=klbwQKm1vEM 0 likes2 downloads3mo agoHugging Face29taqbaylit /hunspell-kab Kabyle Hunspell Dictionary (Cleaned) A cleaned and documented version of M. Belkacem's Imseɣti n tira n teqbaylit (v1.0, MIT), the only comprehensive open-source spell-checker for the Kabyle language (Taqbaylit, ISO 639-1 kab). 📦 Dataset Contents File Description Size kab.dic Cleaned word list with morphological metadata ~685 KB kab.aff Original affixation rules (Belkacem v1.0) ~19 KB cleaning-report.md Full cleaning audit ~6 KB tag-vocabulary.md… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/hunspell-kab.n<1K0 likes2 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.