CoolFace
Datasetpublic

turkish-nlp-suite/temiz-Wiki

A cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo. The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces.

sourceHugging Facecc-by-sa-4.0updated 7mo agoView on Hugging Face
4likes67downloads
Dataset Card

A cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo.

The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces.