datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cm.mgb2arkavidia-final-fekabyle-named-entities
Kabyle Standardized Named Entities Dataset
This is a manually curated parallel corpus in Kabyle complete with semantic English contextual translations and structured Named Entity Recognition (NER) tag assignments.
Dataset Structure
kabyle_standardized: Target entity string conforming to standardized orthographic regulations.
english_translation: High-context semantic meaning, institutional purpose, or micro-topographic geographical breakdowns.
entity_category:… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-named-entities.kabyle-g2p-training-data
Kabyle G2P Training Data
Phonetically-annotated Kabyle (Taqbaylit) text corpus for training Grapheme-to-Phoneme (G2P) models. Generated using the orthography2ipa rule-based phonemizer for Kabyle.
Dataset Overview
Property
Value
Language
Kabyle (kab) — Afro-Asiatic, Berber
Total pairs
59,462
Source
boffire/kabyle-piper-22khz
Phonemizer
orthography2ipa (dev branch)
IPA standard
Narrow transcription with Kabyle-specific allophony
License
CC0… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-g2p-training-data.qwen-1.5b-blind-spotsTechnical Analysis: Blind Spots of Qwen2.5-1.5B
How the Model was Loaded:
The model was loaded in a Google Colab environment using the transformers library with torch_dtype=torch.bfloat16 to fit within the T4 GPU memory limits.
Fine-tuning Strategy:
To fix the identified logical, grammatical, and instruction-following errors, the model should undergo Supervised Fine-Tuning (SFT). This would involve training the model on "Chain-of-Thought" (CoT) datasets where the model is taught to explain its… See the full description on the dataset page: https://huggingface.co/datasets/Taqiiiiiiiii/qwen-1.5b-blind-spots.
