Ba2han/en-tur_corpus
English - Turkish Filtered Corpus This dataset was compiled from multiple English and Turkish sources, heavily filtered, and deduplicated. Processing Applied: Length Filtering: Min characters = 100, Max characters = 2600 Deduplication: Exact deduplication followed by fast heuristic fingerprinting (normalized, whitespace & punctuation removed) Shuffled: Seed = 42 Dataset Composition: Dataset Source Original Rows (Loaded) Final Rows (After… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/en-tur_corpus.
This repository belongs to Ba2han on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
