CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taqbaylit /libretranslate-en-kab-suggestions Kabyle Suggestions Dataset This dataset contains English-to-Kabyle translation suggestions submmitted by users using LibreTranslate, designed to support the development and evaluation of machine translation tools for the Kabyle language. texttranslationn<1K0 likes1.7k downloads4mo agoHugging Face02bh2821 /TaQAQ-TreeD-1mil TaQAQ Series The TaQAQ‑TreeD series is a high‑fidelity, synthetic representation of the NYSE trade and quotes data stream. Designed as part of the broader TaQAQ series, this collection captures one trillion consecutive trading events—each comprising detailed trade executions within a fully simulated market environment that adheres to the same temporal granularity, message formats, and micro‑structure dynamics as real‑world feeds. DATA : T1mil ASOF : 251004 MODAL : Nonfused FORMAT:… See the full description on the dataset page: https://huggingface.co/datasets/bh2821/TaQAQ-TreeD-1mil.0 likes1.5k downloads7mo agoHugging Face03Taquito07 /plant_classification_v11image1K<n<10K0 likes613 downloads3y agoHugging Face04nerusskikh /taqpol_insilico_dms Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: Yulia E. Tomilova, Nikolai E. Russkikh, Igor M. Yi, Elizaveta V. Shaburova, Viktor N. Tomilov, Galina B. Pyrinova, Svetlana O. Brezhneva, Olga S. Tikhonyuk, Nadezhda S. Gololobova, Dmitriy V. Popichenko, Maxim O. Arkhipov, Leonid O. Bryzgalov, Evgeny V. Brenner… See the full description on the dataset page: https://huggingface.co/datasets/nerusskikh/taqpol_insilico_dms.tabular10M<n<100M0 likes427 downloads2y agoHugging Face05taqacuct /kab-ocr-dataset OCR Kabyle (Latin + caractères spéciaux) Dataset synthétique pour entraîner un modèle OCR sur le kabyle. Structure prévue images/ : contiendra les images générées labels.tsv : liste des couples (image → texte) charset.txt : alphabet utilisé image10K<n<100K0 likes189 downloads1y agoHugging Face06taqbaylit /adlis-pdfsdocumentn<1K1 likes142 downloads5mo agoHugging Face07taqbaylit /common-voice-scripted-speech-kab-26-huge Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned) Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips. Source Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12) Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective) License: CC0-1.0 Generated: 2026-07-12 Cleaning Pipeline Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.audio100K<n<1M0 likes91 downloads3mo agoHugging Face08taqbaylit /kabyle-verbs Kabyle Verbs — Kabyle Verb Conjugation Kabyle verb conjugation dataset — 6,198 verbs, ~344,000 conjugated forms, covering aorist, preterite, imperative, participles, and intensive forms. Data source: amyag.com, work by Kamal Nait Zerrad. Summary Property Value Language Kabyle (taqbaylit) Verbs 6,198 Total conjugated forms 344,745 Unique forms 214,276 Grammatical tenses 11 (aorist, preterite, negative preterite, imperative, intensive aorist… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-verbs.text100K<n<1M0 likes86 downloads3mo agoHugging Face09taqwa92 /cm.trial Dataset Card for Common Voice Corpus 11.0 Dataset Summary The Common Voice dataset consists of a unique MP3 and corresponding text file. Many of the 24210 recorded hours in the dataset also include demographic metadata like age, sex, and accent that can help improve the accuracy of speech recognition engines. The dataset currently consists of 16413 validated hours in 100 languages, but more voices and languages are always added. Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/taqwa92/cm.trial.tabularautomatic-speech-recognition10K<n<100K0 likes55 downloads4y agoHugging Face10taqbaylit /Kabyle_Road_Traffic_Code Kabyle-English Road Traffic Code Dataset A bilingual parallel corpus of 102 road traffic signs and regulations in English and Kabyle (Taqbaylit), an Amazigh language spoken in Algeria. Categories Dangers (Imihiten): Warning signs (39 entries) Prohibitions (Tigedlin): Prohibitory signs (35 entries) Obligations (Timariwin): Mandatory signs (16 entries) End of Restrictions: End of regulation signs (12 entries) Splits Split Size Train 62… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/Kabyle_Road_Traffic_Code.tabularn<1K0 likes41 downloads5mo agoHugging Face11taqbaylit /f5tts-kabyle-dataset F5-TTS Kabyle Dataset Clean, deduplicated audio-text dataset for Kabyle (Taqbaylit / Tamaziɣt) TTS fine-tuning with F5-TTS. Statistics Metric Value Total clips 59,462 Total duration 41.30 hours Sample rate 24 kHz mono Avg clip length 2.50s Min clip length 1.00s Max clip length 12.65s Unique phrases 59,462 (0% duplicates) Unique characters 112 Sources Tatoeba (67.8%) + Common Voice 26 tiny (32.2%) Source Datasets… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/f5tts-kabyle-dataset.text10K<n<100K0 likes36 downloads3mo agoHugging Face12taqbaylit /bejaia Béjaïa University Theses Dataset (Taqbaylit / Kabyle) Description This dataset contains 647 theses from the Université Abderrahmane Mira de Béjaïa DSpace institutional repository in Algeria: Folder Count Level master/ 640 Master theses (Mémoires de Master) magister/ 7 Magister theses (Mémoires de Magistère) All documents are in PDF format and were collected from the university's public open-access archive. The theses focus on Kabyle (Taqbaylit)… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/bejaia.documentn<1K0 likes30 downloads4mo agoHugging Face13taqbaylit /weblate-kabyle0 likes30 downloads12d agoHugging Face14Taquito07 /plant_classification_v2image1K<n<10K3 likes26 downloads3y agoHugging Face15taqwa92 /mg.trial4audio1K<n<10K0 likes25 downloads4y agoHugging Face16taqbaylit /tatoeba-kabyle-mono-cleaned tatoeba-kabyle-mono-cleaned Cleaned and quality-assessed monolingual Kabyle corpus extracted from Tatoeba. Summary This dataset contains sentences from Tatoeba tagged as Kabyle (lang == "kab"), processed through a multi-layer linguistic filtering pipeline combining orthographic normalization, language identification (GlotLID v3 + DistilBERT Kabyle/Tachelhit classifier), code-switching detection (MaskLID), and lexical validation (Kabyle Hunspell dictionary).… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-mono-cleaned.tabular100K<n<1M0 likes24 downloads2mo agoHugging Face17taqwa92 /mg2_dataaudio10K<n<100K0 likes20 downloads4y agoHugging Face18taqwa92 /cm.mgb2tabular10K<n<100K1 likes18 downloads4y agoHugging Face19azrunguraya /kabyle-audio-taqbaylit-languageaudion<1K0 likes15 downloads1y agoHugging Face20taqbaylit /bouira Bouira University Magistère Theses Dataset Description This dataset contains 395 magistère (master's) theses from the Université de Bouira DSpace institutional repository in Algeria. All documents are in PDF format and were from the university's public open-access archive. documentn<1K1 likes15 downloads5mo agoHugging Face21taqiyudinadn /arkavidia-final-fetabular1M<n<10M0 likes14 downloads7mo agoHugging Face22taqbaylit /kabyle-toponyms Algeria French–Kabyle Toponym Corpus A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab. Description This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI). The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.tabulartranslation1K<n<10K0 likes14 downloads4mo agoHugging Face23taqbaylit /tatoeba-en-kab Tatoeba English-Kabyle Parallel Corpus A cleaned and aligned English-Kabyle parallel corpus extracted from Tatoeba, with both direct en↔kab links and indirect kab→fra→en links. Statistics Split Pairs train 240,056 dev 2,449 test 2,449 Total 244,954 Source Tatoeba direct en↔kab links Indirect kab→fra→en links (Kabyle linked to French, French linked to English) Cleaning Pipeline Character standardization: Fixed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-en-kab.texttranslation100K<n<1M0 likes14 downloads3mo agoHugging Face24taqbaylit /ayamun-pdfs230 pdf files from Ayamun. documentn<1K1 likes13 downloads4mo agoHugging Face25taqbaylit /tatoeba-kabyle-audio Tatoeba Kabyle Audio Dataset A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected. Dataset Description This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/tatoeba-kabyle-audio.audioautomatic-speech-recognition10K<n<100K0 likes13 downloads3mo agoHugging Face26taqwa92 /mg21_data0 likes12 downloads4y agoHugging Face27taqbaylit /kabyle-corpustext100K<n<1M0 likes12 downloads4mo agoHugging Face28taqiyudinadn /EOS-Continued-Pretraining-Dataset EOS Continued Pre-Training Dataset (Indonesia) Deskripsi Dataset EOS Continued Pre-Training Dataset adalah korpus teks bahasa Indonesia yang dikurasi untuk proses Continued Pre-Training (CPT) pada Large Language Models (LLM). Tujuan utama dari dataset ini adalah untuk melakukan Domain Adaptation, yaitu meningkatkan kemampuan model dalam memahami konteks, terminologi, dan nuansa pada dua domain strategis di Indonesia: Pengawasan Ruang Digital (PRD) Digital Talent Pool… See the full description on the dataset page: https://huggingface.co/datasets/taqiyudinadn/EOS-Continued-Pretraining-Dataset.texttext-generation100K<n<1M0 likes11 downloads9mo agoHugging Face29taqwa92 /mgb2_data1 likes10 downloads4y agoHugging Face30taqbaylit /kabyle-synth-voice Kabyle Parallel Corpus (OmniVoice × Tatoeba) Corpus parallèle de 997 phrases kabyles avec audio généré par OmniVoice. Statistiques Langue : kabyle (kab) Phrases totales : 997 Nouvelles phrases (ce run) : 987 Durée totale : 1958.4s (32.6 min) Sampling rate : 24000 Hz Source texte : Tatoeba Modèle TTS : k2-fsa/OmniVoice Dernière mise à jour : 2026-05-10T10:49:01.178795 Structure kabyle_corpus_997/ ├── audio/ # Fichiers WAV ├──… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-synth-voice.0 likes10 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.