CoolFace
Datasetpublic

diffutron/DiffutronLM-Pretraining-Corpus

DiffutronLM-Pretraining-Corpus DiffutronLM-Pretraining-Corpus is the comprehensive, filtered Turkish text dataset used during the Continual Pre-training (CPT) phase of the Diffutron language models. The primary goal of this dataset was to align the cross-lingual representations of a multilingual base encoder (jhu-clsp/mmBERT-base) with the agglutinative complexity and morphological nuances of the Turkish language, without inducing catastrophic forgetting. 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/diffutron/DiffutronLM-Pretraining-Corpus.

sourceHugging Faceupdated 6mo agoView on Hugging Face
2likes51downloads

diffutron/DiffutronLM-Pretraining-Corpus · main · files are served by the source, never re-hosted here