CoolFace
Datasetpublic

orionweller/mmBERT-pretraining-data-chunk1

mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk1.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes5.3kdownloads
settings

This repository belongs to orionweller on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemmBERT-pretraining-data-chunk1
visibilitypublic
licencemit
gatedno
ownerorionweller
Account settings
orionweller/mmBERT-pretraining-data-chunk1 · CoolFace