Imsidag-community/nllb_en_kab
NLLB English - Kabyle Dataset This dataset contains parallel sentences in English and Kabyle, cleaned and filtered using the GlotLid model. The dataset is derived from the OPUS-NLLB corpus and has been processed to ensure high-quality sentence pairs. Dataset Structure nllb_en_kab.parquet: A Parquet file containing the cleaned English-Kabyle sentence pairs. Dataset Statistics Total Sentence Pairs: 2,484,297 English Sentences: 2,484,297 Kabyle… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/nllb_en_kab.
224
