Ba2han/en-tur_corpus
English - Turkish Filtered Corpus This dataset was compiled from multiple English and Turkish sources, heavily filtered, and deduplicated. Processing Applied: Length Filtering: Min characters = 100, Max characters = 2600 Deduplication: Exact deduplication followed by fast heuristic fingerprinting (normalized, whitespace & punctuation removed) Shuffled: Seed = 42 Dataset Composition: Dataset Source Original Rows (Loaded) Final Rows (After… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/en-tur_corpus.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face