datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midtraining_mix_modernbert_filtered_documentsmailroom-modernbert-training
mailroom-modernbert-training
Cleaned + prepared hierarchical-classification training set for the
ModernBERT-base ingest fast-path — the fine-tuning surface of the
mailroom-ml synthetic-data layer.
The layer (end to end)
Lucius-Morningstar/mailroom-dataset corpus (GT labels, canonical v9)
│ (working copy, pinned)
▼
THIS REPO (mailroom-modernbert-training) @ pinned revision
│ training/train_modernbert.py (--data <repo>)
▼… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/mailroom-modernbert-training.ni-ood-dataset-20250131-modernbert-train-kmeans-dim128-20250312dolma3-hq-2M-modernbert
Dolma3 High-Quality 2M (ModernBERT Filtered)
A curated subset of 2 million high-quality text samples from allenai/dolma3_dolmino_mix-100B-1125, filtered to fit within ModernBERT's 8192 token context window.
Dataset Description
This dataset is designed for pretraining diffusion language models based on ModernBERT. Each sample has been:
Source filtered: Only from ingredient1-common_crawl-high-quality folders (highest quality web text)
Length filtered: Minimum 200… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/dolma3-hq-2M-modernbert.wikitext-tags-modernbertdata_ablation_full59K-modernbert-split-kmeans-dim768-20250218Russian-toxic-modernbert
Tokenized Russian toxic text
Tokenized version of Mnwa/russian-toxic dataset with modernbert base model tokenizer
fineweb-10b-512-modernbert
FineWeb-Edu — ModernBERT continuous packed chunks
Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files).
Tokenizer: answerdotai/ModernBERT-large. No truncation or padding.
Each nonempty document contributes CLS (50281), document IDs, SEP (50282).
The concatenated stream is split into 512-token rows. Documents
may span chunks; a chunk need not begin with CLS or end with SEP.
Original sorted-file and row order is preserved throughout processing.… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-10b-512-modernbert.ni-20-clustered-fulltext-modernbert-sweep-20250107ni-ood-dataset-20250131-modernbert-train-kmeans-dim768-20250318ni-20-clustered-fulltext-modernbert-sweep-20250107-modernbert-split-kmeans-dim768-20250130paws-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the google-research-datasets/paws dataset
This is the google-research-datasets/paws dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/paws-gte-modernbert-pooled.chonkiepedia-modernbert-tokenizedpubmedqa-query-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the qiaojin/PubMedQA dataset
This is the qiaojin/PubMedQA dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to the model’s maximum token… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/pubmedqa-query-gte-modernbert-pooled.ni-ood-dataset-10p-20250127-modernbert-kmeans-dim128-20250128SlimPajama-6B-modernbert-split-kmeans-dim768-20250316msmarco-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/msmarco-corpus dataset
This is the sentence-transformers/msmarco-corpus dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/msmarco-gte-modernbert-pooled.final_train_test_split_modernbert_largedata_ablation_full59K-modernbert-split-kmeans-dim768-20250321clt-eval-modernbert-tokenizeddolly-15k-clustered-fulltext-modernbert-sweep-20250106ni-unique-20-tasks-modernbert-dbscan-dim128-silscore0.48760950565338135-20250123modernbert_encoder_sp_seq_512_csedm_fold1eval-gliner2-modernbert_pasteproof-uni-20260621modernbert_encoder_sp_seq_512_dbe22kt_fold1ni-ood-dataset-10p-20250127-modernbert-split-kmeans-dim128-20250128databricks-dolly-15k-modernbert-train-kmeans-dim768-20250723ModernBERT-512-Combined-v3ni-unique-20-tasks-modernbert-dbscan-dim64-silscore-100-20250120ni-unique-100-tasks-modernbert-split-kmeans-dim768-20250310
