CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taqbaylit /libretranslate-en-kab-suggestions Kabyle Suggestions Dataset This dataset contains English-to-Kabyle translation suggestions submmitted by users using LibreTranslate, designed to support the development and evaluation of machine translation tools for the Kabyle language. texttranslationn<1K0 likes1.7k downloads4mo agoHugging Face02nerusskikh /taqpol_insilico_dms Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: Yulia E. Tomilova, Nikolai E. Russkikh, Igor M. Yi, Elizaveta V. Shaburova, Viktor N. Tomilov, Galina B. Pyrinova, Svetlana O. Brezhneva, Olga S. Tikhonyuk, Nadezhda S. Gololobova, Dmitriy V. Popichenko, Maxim O. Arkhipov, Leonid O. Bryzgalov, Evgeny V. Brenner… See the full description on the dataset page: https://huggingface.co/datasets/nerusskikh/taqpol_insilico_dms.tabular10M<n<100M0 likes431 downloads2y agoHugging Face03taqbaylit /common-voice-scripted-speech-kab-26-huge Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned) Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips. Source Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12) Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective) License: CC0-1.0 Generated: 2026-07-12 Cleaning Pipeline Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.audio100K<n<1M0 likes336 downloads3mo agoHugging Face04taqbaylit /kabyle-verbs Kabyle Verbs — Kabyle Verb Conjugation Kabyle verb conjugation dataset — 6,198 verbs, ~344,000 conjugated forms, covering aorist, preterite, imperative, participles, and intensive forms. Data source: amyag.com, work by Kamal Nait Zerrad. Summary Property Value Language Kabyle (taqbaylit) Verbs 6,198 Total conjugated forms 344,745 Unique forms 214,276 Grammatical tenses 11 (aorist, preterite, negative preterite, imperative, intensive aorist… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-verbs.text100K<n<1M0 likes86 downloads3mo agoHugging Face05taqwa92 /cm.trial Dataset Card for Common Voice Corpus 11.0 Dataset Summary The Common Voice dataset consists of a unique MP3 and corresponding text file. Many of the 24210 recorded hours in the dataset also include demographic metadata like age, sex, and accent that can help improve the accuracy of speech recognition engines. The dataset currently consists of 16413 validated hours in 100 languages, but more voices and languages are always added. Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/taqwa92/cm.trial.tabularautomatic-speech-recognition10K<n<100K0 likes54 downloads4y agoHugging Face06taqbaylit /Kabyle_Road_Traffic_Code Kabyle-English Road Traffic Code Dataset A bilingual parallel corpus of 102 road traffic signs and regulations in English and Kabyle (Taqbaylit), an Amazigh language spoken in Algeria. Categories Dangers (Imihiten): Warning signs (39 entries) Prohibitions (Tigedlin): Prohibitory signs (35 entries) Obligations (Timariwin): Mandatory signs (16 entries) End of Restrictions: End of regulation signs (12 entries) Splits Split Size Train 62… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/Kabyle_Road_Traffic_Code.tabularn<1K0 likes43 downloads5mo agoHugging Face07taqbaylit /f5tts-kabyle-dataset F5-TTS Kabyle Dataset Clean, deduplicated audio-text dataset for Kabyle (Taqbaylit / Tamaziɣt) TTS fine-tuning with F5-TTS. Statistics Metric Value Total clips 59,462 Total duration 41.30 hours Sample rate 24 kHz mono Avg clip length 2.50s Min clip length 1.00s Max clip length 12.65s Unique phrases 59,462 (0% duplicates) Unique characters 112 Sources Tatoeba (67.8%) + Common Voice 26 tiny (32.2%) Source Datasets… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/f5tts-kabyle-dataset.text10K<n<100K0 likes36 downloads3mo agoHugging Face08taqbaylit /tatoeba-kabyle-mono-cleaned tatoeba-kabyle-mono-cleaned Cleaned and quality-assessed monolingual Kabyle corpus extracted from Tatoeba. Summary This dataset contains sentences from Tatoeba tagged as Kabyle (lang == "kab"), processed through a multi-layer linguistic filtering pipeline combining orthographic normalization, language identification (GlotLID v3 + DistilBERT Kabyle/Tachelhit classifier), code-switching detection (MaskLID), and lexical validation (Kabyle Hunspell dictionary).… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-mono-cleaned.tabular100K<n<1M0 likes24 downloads2mo agoHugging Face09taqwa92 /mg.trial4audio1K<n<10K0 likes23 downloads4y agoHugging Face10taqwa92 /mg2_dataaudio10K<n<100K0 likes20 downloads4y agoHugging Face11taqwa92 /cm.mgb2tabular10K<n<100K1 likes18 downloads4y agoHugging Face12taqbaylit /kabyle-toponyms Algeria French–Kabyle Toponym Corpus A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab. Description This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI). The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.tabulartranslation1K<n<10K0 likes15 downloads4mo agoHugging Face13taqbaylit /tatoeba-en-kab Tatoeba English-Kabyle Parallel Corpus A cleaned and aligned English-Kabyle parallel corpus extracted from Tatoeba, with both direct en↔kab links and indirect kab→fra→en links. Statistics Split Pairs train 240,056 dev 2,449 test 2,449 Total 244,954 Source Tatoeba direct en↔kab links Indirect kab→fra→en links (Kabyle linked to French, French linked to English) Cleaning Pipeline Character standardization: Fixed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-en-kab.texttranslation100K<n<1M0 likes14 downloads3mo agoHugging Face14taqbaylit /ayamun-pdfs230 pdf files from Ayamun. documentn<1K1 likes13 downloads4mo agoHugging Face15taqbaylit /tatoeba-kabyle-audio Tatoeba Kabyle Audio Dataset A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected. Dataset Description This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.audioautomatic-speech-recognition10K<n<100K0 likes13 downloads3mo agoHugging Face16taqiyudinadn /arkavidia-final-fetabular1M<n<10M0 likes12 downloads7mo agoHugging Face17taqbaylit /kabyle-corpustext100K<n<1M0 likes12 downloads4mo agoHugging Face18taqbaylit /kabyle-english-translatewiki English-Kabyle Parallel Corpus A clean, deduplicated parallel corpus of English → Kabyle (Taqbaylit) translations extracted from the translatewiki.net bulk dump (2026-01-01). Dataset Summary Attribute Value Language pair English (en) → Kabyle (kab) Total unique pairs 8,871 Source translatewiki.net License CC BY 3.0 Domain Software localization, UI strings, documentation Dataset Structure { "translation": { "en": "Hello"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-translatewiki.texttranslation1K<n<10K0 likes10 downloads4mo agoHugging Face19taqbaylit /kab-en-toponyms-sentences English-Kabyle Parallel Corpus for Machine Translation This dataset contains 32,024 grammatically flawless parallel sentence pairs mapping English to literary Kabyle (Taqbaylit kab). This corpus was synthesized using a linguistically-informed morphosyntactic rule engine paired with clean OpenStreetMap toponym registries from boffire/kabyle-toponyms. It handles complex phonetic mutations natively, making it a state-of-the-art bootstrapping asset for fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kab-en-toponyms-sentences.text10K<n<100K0 likes10 downloads4mo agoHugging Face20taqbaylit /kabyle-named-entities Kabyle Standardized Named Entities Dataset This is a manually curated parallel corpus in Kabyle complete with semantic English contextual translations and structured Named Entity Recognition (NER) tag assignments. Dataset Structure kabyle_standardized: Target entity string conforming to standardized orthographic regulations. english_translation: High-context semantic meaning, institutional purpose, or micro-topographic geographical breakdowns. entity_category:… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-named-entities.texttoken-classificationn<1K0 likes10 downloads4mo agoHugging Face21taqiyudinadn /EOS-Continued-Pretraining-Dataset EOS Continued Pre-Training Dataset (Indonesia) Deskripsi Dataset EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia yang dikurasi untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM). Tujuan utama dari dataset ini adalah untuk melakukan Domain Adaptation, yaitu meningkatkan kemampuan model dalam memahami konteks, terminologi, dan nuansa pada dua domain strategis di Indonesia: Pengawasan Ruang Digital (PRD) Digital Talent Pool… See the full description on the dataset page: https://huggingface.co/datasets/taqiyudinadn/EOS-Continued-Pretraining-Dataset.texttext-generation100K<n<1M0 likes9 downloads9mo agoHugging Face22taqbaylit /nllb_en_kab NLLB English-Kabyle Parallel Corpus (Filtered & Cleaned) Parallel English–Kabyle sentence pairs derived from the OPUS-NLLB corpus, filtered with GlotLid v3 and cleaned through a multi-stage Kabyle-specific pipeline. Dataset Structure nllb_en_kab.parquet: Parquet file with two columns: english: English sentence kabyle: Kabyle sentence Statistics Metric Count Total sentence pairs 2,786,012 Non-null English 2,786,012 Non-null Kabyle… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/nllb_en_kab.texttranslation1M<n<10M0 likes9 downloads3mo agoHugging Face23taqbaylit /timucuha-kabyle-tales Timucuha Trilingual Corpus A parallel corpus of Kabyle (Tamazight) folk tales with French and English translations. Source The original Kabyle tales were collected and digitized by the Association Culturelle Numidya. This dataset is derived from their Timucuha project, which preserves and promotes Kabyle oral tradition. Website: https://timucuha.numidya.net/ Organization: Association Culturelle Numidya Dataset Description Property Value… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/timucuha-kabyle-tales.texttranslationn<1K0 likes8 downloads4mo agoHugging Face24taqbaylit /rradyu-tis-snat Rradyu Tis Snat — Kabyle Podcasts from Radio Algérie Chaîne 2 Status: work in progress. This README is a first draft with placeholders (marked TODO) to fill in as the dataset grows. Metadata above (license, size_categories) will need updating as the collection is built out. Dataset Description This dataset is a collection of Kabyle-language ("Taqbaylit") audio podcasts from Radio Algérie Chaîne 2 (podcast.radioalgerie.dz), the Algerian public radio channel… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/rradyu-tis-snat.audioautomatic-speech-recognitionn<1K0 likes5 downloads2mo agoHugging Face25TA-LLM /TA_Quote_Codetext1K<n<10K0 likes4 downloads3y agoHugging Face26Bixente-san /humour-taquin-francaistextn<1K0 likes4 downloads8mo agoHugging Face27taqbaylit /kabyle-english-TM Kabyle–English Translation Memory A bilingual translation memory containing 121,725 sentence pairs in Kabyle (kab) and English (en), built from open-source software localisation data aggregated through an automated pipeline. Dataset structure Each record contains the following fields: Field Type Description source string Source segment (English) source_lang string Always "en" target string Target segment (Kabyle) target_lang string Always "kab"… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-english-TM.texttranslation100K<n<1M0 likes4 downloads2mo agoHugging Face28taqbaylit /kabyle-g2p-training-data Kabyle G2P Training Data Phonetically-annotated Kabyle (Taqbaylit) text corpus for training Grapheme-to-Phoneme (G2P) models. Generated using the orthography2ipa rule-based phonemizer for Kabyle. Dataset Overview Property Value Language Kabyle (kab) — Afro-Asiatic, Berber Total pairs 59,462 Source boffire/kabyle-piper-22khz Phonemizer orthography2ipa (dev branch) IPA standard Narrow transcription with Kabyle-specific allophony License CC0… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-g2p-training-data.texttext-to-speech10K<n<100K1 likes4 downloads2mo agoHugging Face29taqwa92 /cm.trial1text1K<n<10K0 likes3 downloads4y agoHugging Face30Taqiiiiiiiii /qwen-1.5b-blind-spotsTechnical Analysis: Blind Spots of Qwen2.5-1.5B How the Model was Loaded: The model was loaded in a Google Colab environment using the transformers library with torch_dtype=torch.bfloat16 to fit within the T4 GPU memory limits. Fine-tuning Strategy: To fix the identified logical, grammatical, and instruction-following errors, the model should undergo Supervised Fine-Tuning (SFT). This would involve training the model on "Chain-of-Thought" (CoT) datasets where the model is taught to explain its… See the full description on the dataset page: https://huggingface.co/datasets/Taqiiiiiiiii/qwen-1.5b-blind-spots.textn<1K0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.