shiima/vejin-Dataset-Normalization
Kurdish Books Dataset (Preprocessed) Dataset Description This dataset contains 18,565 Kurdish books with asosoft preprocessing applied to the content field. The dataset was created from an Excel file and includes book metadata along with preprocessed text content. Languages Central Kurdish (ckb) Kurdish (ku) Dataset Structure The dataset contains the following columns: author book title url content Data Processing… See the full description on the dataset page: https://huggingface.co/datasets/shiima/vejin-Dataset-Normalization.
This repository belongs to shiima on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
