CoolFace
Datasetpublicgated

shiima/kurdish-unified-corpus

Unified Kurdish Corpus Dataset Description This dataset aggregates 758,166 Kurdish text samples from multiple high-quality sources. All text has been preprocessed using asosoft. Languages Central Kurdish (ckb) Kurdish (ku) Dataset Structure Columns text: Preprocessed text content (asosoft applied) base_dataset: Source dataset name url: Source URL (NULL if not available) word_count: Number of words (space-separated)… See the full description on the dataset page: https://huggingface.co/datasets/shiima/kurdish-unified-corpus.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes5downloads
settings

This repository belongs to shiima on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namekurdish-unified-corpus
visibilitypublic
licenceapache-2.0
gatedyes
ownershiima
Account settings
shiima/kurdish-unified-corpus · CoolFace