CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01geodesic-research /midtraining_mix_modernbert_filtered_documentstext1M<n<10M0 likes307 downloads10mo agoHugging Face02Lucius-Morningstar /mailroom-modernbert-training mailroom-modernbert-training Cleaned + prepared hierarchical-classification training set for the ModernBERT-base ingest fast-path — the fine-tuning surface of the mailroom-ml synthetic-data layer. The layer (end to end) Lucius-Morningstar/mailroom-dataset corpus (GT labels, canonical v9) │ (working copy, pinned) ▼ THIS REPO (mailroom-modernbert-training) @ pinned revision │ training/train_modernbert.py (--data <repo>) ▼… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/mailroom-modernbert-training.tabulartext-classification1K<n<10K1 likes165 downloads19h agoHugging Face03rchu233 /ni-ood-dataset-20250131-modernbert-train-kmeans-dim128-20250312tabular1M<n<10M0 likes146 downloads2y agoHugging Face04Ayushnangia /dolma3-hq-2M-modernbert Dolma3 High-Quality 2M (ModernBERT Filtered) A curated subset of 2 million high-quality text samples from allenai/dolma3_dolmino_mix-100B-1125, filtered to fit within ModernBERT's 8192 token context window. Dataset Description This dataset is designed for pretraining diffusion language models based on ModernBERT. Each sample has been: Source filtered: Only from ingredient1-common_crawl-high-quality folders (highest quality web text) Length filtered: Minimum 200… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/dolma3-hq-2M-modernbert.texttext-generation1M<n<10M0 likes146 downloads8mo agoHugging Face05hriaz /wikitext-tags-modernberttext1M<n<10M0 likes124 downloads1y agoHugging Face06albertge /data_ablation_full59K-modernbert-split-kmeans-dim768-20250218tabular10K<n<100K0 likes120 downloads2y agoHugging Face07Mnwa /Russian-toxic-modernbert Tokenized Russian toxic text Tokenized version of Mnwa/russian-toxic dataset with modernbert base model tokenizer text-classification100K<n<1M0 likes94 downloads2y agoHugging Face08dme5245 /fineweb-10b-512-modernbert FineWeb-Edu — ModernBERT continuous packed chunks Source: HuggingFaceFW/fineweb-edu, sample-10BT (or the supplied local Parquet files). Tokenizer: answerdotai/ModernBERT-large. No truncation or padding. Each nonempty document contributes CLS (50281), document IDs, SEP (50282). The concatenated stream is split into 512-token rows. Documents may span chunks; a chunk need not begin with CLS or end with SEP. Original sorted-file and row order is preserved throughout processing.… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-10b-512-modernbert.10M<n<100M0 likes55 downloads5d agoHugging Face09albertge /ni-20-clustered-fulltext-modernbert-sweep-20250107tabular10K<n<100K0 likes54 downloads2y agoHugging Face10NamburiSrinath /ni-ood-dataset-20250131-modernbert-train-kmeans-dim768-20250318tabular1M<n<10M0 likes49 downloads2y agoHugging Face11rchu233 /ni-20-clustered-fulltext-modernbert-sweep-20250107-modernbert-split-kmeans-dim768-20250130tabular10K<n<100K0 likes47 downloads2y agoHugging Face12stephantulkens /paws-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the google-research-datasets/paws dataset This is the google-research-datasets/paws dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/paws-gte-modernbert-pooled.text1M<n<10M0 likes42 downloads11mo agoHugging Face13feyninc /chonkiepedia-modernbert-tokenized1M<n<10M0 likes41 downloads1y agoHugging Face14stephantulkens /pubmedqa-query-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the qiaojin/PubMedQA dataset This is the qiaojin/PubMedQA dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to the model’s maximum token… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/pubmedqa-query-gte-modernbert-pooled.text100K<n<1M0 likes41 downloads11mo agoHugging Face15rchu233 /ni-ood-dataset-10p-20250127-modernbert-kmeans-dim128-20250128tabular100K<n<1M0 likes34 downloads2y agoHugging Face16albertge /SlimPajama-6B-modernbert-split-kmeans-dim768-20250316tabular1M<n<10M0 likes33 downloads2y agoHugging Face17stephantulkens /msmarco-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/msmarco-corpus dataset This is the sentence-transformers/msmarco-corpus dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/msmarco-gte-modernbert-pooled.text1M<n<10M0 likes33 downloads11mo agoHugging Face18waleedk7-fyp /final_train_test_split_modernbert_large0 likes32 downloads5mo agoHugging Face19albertge /data_ablation_full59K-modernbert-split-kmeans-dim768-20250321tabular10K<n<100K0 likes31 downloads2y agoHugging Face20bluelightai-dev /clt-eval-modernbert-tokenizedtabular100K<n<1M0 likes31 downloads7mo agoHugging Face21albertge /dolly-15k-clustered-fulltext-modernbert-sweep-20250106tabular10K<n<100K0 likes30 downloads2y agoHugging Face22NamburiSrinath /ni-unique-20-tasks-modernbert-dbscan-dim128-silscore0.48760950565338135-20250123text10K<n<100K0 likes26 downloads2y agoHugging Face23Unggi /modernbert_encoder_sp_seq_512_csedm_fold1tabularn<1K0 likes25 downloads2y agoHugging Face24AITeamUIT /eval-gliner2-modernbert_pasteproof-uni-202606210 likes25 downloads3mo agoHugging Face25Unggi /modernbert_encoder_sp_seq_512_dbe22kt_fold1tabularn<1K0 likes24 downloads2y agoHugging Face26rchu233 /ni-ood-dataset-10p-20250127-modernbert-split-kmeans-dim128-20250128tabular100K<n<1M0 likes21 downloads2y agoHugging Face27albertge /databricks-dolly-15k-modernbert-train-kmeans-dim768-20250723tabular10K<n<100K0 likes21 downloads1y agoHugging Face28JamesResearch1216 /ModernBERT-512-Combined-v310M<n<100M1 likes21 downloads3mo agoHugging Face29NamburiSrinath /ni-unique-20-tasks-modernbert-dbscan-dim64-silscore-100-20250120text10K<n<100K0 likes20 downloads2y agoHugging Face30rchu233 /ni-unique-100-tasks-modernbert-split-kmeans-dim768-20250310tabular100K<n<1M0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.