CoolFace
Datasetpublic

orionweller/mmBERT-pretraining-data-chunk3

mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk3.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes1.3kdownloads

orionweller/mmBERT-pretraining-data-chunk3 · main · files are served by the source, never re-hosted here

orionweller/mmBERT-pretraining-data-chunk3 · CoolFace