CoolFace
Datasetpublic

Ba2han/en-tur_corpus

English - Turkish Filtered Corpus This dataset was compiled from multiple English and Turkish sources, heavily filtered, and deduplicated. Processing Applied: Length Filtering: Min characters = 100, Max characters = 2600 Deduplication: Exact deduplication followed by fast heuristic fingerprinting (normalized, whitespace & punctuation removed) Shuffled: Seed = 42 Dataset Composition: Dataset Source Original Rows (Loaded) Final Rows (After… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/en-tur_corpus.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes229downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Ba2han/en-tur_corpus · CoolFace