CoolFace
Datasetpublic

Ba2han/finepdfs-long

HuggingFaceFW/finepdfs long filtered Turkish texts Source: HuggingFaceFW/finepdfs (config: tur_Latn). Rows contain 4,000–15,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected rows: 140,166. Generated by process_hf_dataset.py. See summary.json for counts and thresholds.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes34downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Ba2han/finepdfs-long · CoolFace