CoolFace
Datasetpublic

Ba2han/finepdfs-long

HuggingFaceFW/finepdfs long filtered Turkish texts Source: HuggingFaceFW/finepdfs (config: tur_Latn). Rows contain 4,000–15,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected rows: 140,166. Generated by process_hf_dataset.py. See summary.json for counts and thresholds.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes47downloads
2 commits on main
9b740c32mo ago

Replace dataset with iteration-5 long-text filtering output

Ba2han
b5c27f62mo ago

initial commit

Ba2han