mmbert
Datasets
All datasets matching “mmbert”mmBERT-pretrain-p2-fineweb2-remaining
mmBERT Pre-training Data P2
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P2 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p2-fineweb2-remaining.mmBERT-decay-data
MMBERT Decay Phase Data
Phase 3 of 3: Annealed language learning decay phase (100B tokens) with massive multilingual expansion to 1833 languages.
📊 Data Composition
NOTE: there are multiple decay data mixtures: this mixture described below is the Decay-Cont mixture. However, the data in this repository is the Decay-Eng. If you are interested in the others, please let me know so I can prioritize it.
Data Source
Tokens (B)
Percentage
Description
FineWeb2… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/mmBERT-decay-data.mmBERT-midtraining-data
mmBERT Mid-training Data
Phase 2 of 3: High-quality mid-training data mixture (600B tokens) with context extension to 8192 tokens.
This dataset contains the mid-training phase data used to train all mmBERT encoder models. This phase focuses on higher quality data sources and extends the context length from 1024 to 8192 tokens. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.
📊 Data Composition
Data Source… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/mmBERT-midtraining-data.mmBERT-pretrain-p3-others
mmBERT Pre-training Data P3
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P3 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p3-others.mmBERT-pretrain-p1-fineweb2-langs
mmBERT Pre-training Data P1
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P1 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p1-fineweb2-langs.mmBERT-pretraining-data-chunk1
mmBERT Training Data (Ready-to-Use)
Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training.
This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk1.
