CoolFace
Datasetpublic

Ba2han/en-tur_corpus

English - Turkish Filtered Corpus This dataset was compiled from multiple English and Turkish sources, heavily filtered, and deduplicated. Processing Applied: Length Filtering: Min characters = 100, Max characters = 2600 Deduplication: Exact deduplication followed by fast heuristic fingerprinting (normalized, whitespace & punctuation removed) Shuffled: Seed = 42 Dataset Composition: Dataset Source Original Rows (Loaded) Final Rows (After… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/en-tur_corpus.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes229downloads
settings

This repository belongs to Ba2han on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameen-tur_corpus
visibilitypublic
licencenot set
gatedno
ownerBa2han
Account settings
Ba2han/en-tur_corpus · CoolFace