CoolFace
Datasetpublic

Imsidag-community/nllb_en_kab

NLLB English - Kabyle Dataset This dataset contains parallel sentences in English and Kabyle, cleaned and filtered using the GlotLid model. The dataset is derived from the OPUS-NLLB corpus and has been processed to ensure high-quality sentence pairs. Dataset Structure nllb_en_kab.parquet: A Parquet file containing the cleaned English-Kabyle sentence pairs. Dataset Statistics Total Sentence Pairs: 2,484,297 English Sentences: 2,484,297 Kabyle… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/nllb_en_kab.

sourceHugging Faceotherupdated 11mo agoView on Hugging Face
2likes24downloads
filenllb_en_kab.parquet166.1 MBdownload

Imsidag-community/nllb_en_kab · main · files are served by the source, never re-hosted here