turkish-nlp-suite/temiz-Wiki
A cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo. The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces.
minor
Update README.md
Update README.md
Upload data/train/wiki_3.jsonl with huggingface_hub
Upload data/train/wiki_2.jsonl with huggingface_hub
Upload data/train/wiki_1.jsonl with huggingface_hub
Upload data/train/wiki_0.jsonl with huggingface_hub
Upload data/train/wiki_3.jsonl with huggingface_hub
Upload data/train/wiki_2.jsonl with huggingface_hub
Upload data/train/wiki_1.jsonl with huggingface_hub
Upload data/train/wiki_0.jsonl with huggingface_hub
initial commit
