datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
contrastive-stubsmmarco-contrastive
mMARCO-contrastive
The dataset is a modification of mMARCO focusing on French and English parts. The aim is to train a
bi-encoder model using all hard negatives from the database. Instead of having a query/positive/negative triplet, we pair all negatives with a query and a
positive. However, it's worth noting that there are many false negatives in the dataset. This isn't a big issue with a triplet view because false negatives
are much fewer in number, but it's more significant with… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/mmarco-contrastive.contrastive-probing-maccontrastive-pretraining
Contrastive Pretraining
Per-language query/document pairs produced by the retrieval-common-crawl pipeline.
Each config corresponds to a single language or source with identical LightOn-style schema.
Config overview
Configs available are: fw-edu, fw2-arb_Arab, fw2-ces_Latn, fw2-cmn_Hani, fw2-dan_Latn, fw2-deu_Latn, fw2-ell_Grek, fw2-fas_Arab, fw2-fra_Latn, fw2-hun_Latn, fw2-ind_Latn, fw2-ita_Latn, fw2-jpn_Jpan, fw2-nld_Latn, fw2-pol_Latn, fw2-por_Latn, fw2-rus_Cyrl… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/contrastive-pretraining.laion_audio_contrastive
⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models.
Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive
📜 Citation
If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/laion_audio_contrastive.BidirLM-Contrastive
BidirLM-Contrastive
The contrastive training dataset used to train BidirLM Embedding models. It contains 10,110,219 query-document pairs from 79 base datasets, split into 203 subdatasets by language or type (~13 GB), covering three sources: Nemotron, KaLM, and parallel/other data. This dataset is described in the paper: BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs.
If you use this dataset in your research or applications, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/BidirLM-Contrastive.halvest-contrastive
HALvest-Contrastive
Contrastive triplets Harvested from HAL
Citation
@misc{kulumba2026doesauthorshipsignalemerge,
title={Where Does Authorship Signal Emerge in Encoder-Based Language Models?},
author={Francis Kulumba and Guillaume Vimont and Laurent Romary and Florian Cafiero},
year={2026},
eprint={2605.19908},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2605.19908},
}… See the full description on the dataset page: https://huggingface.co/datasets/almanach/halvest-contrastive.mscoco_contrastive
⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models.
Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive
📜 Citation
If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/mscoco_contrastive.halvest-contrastive-oldcontrastive-training-base-bundles
base_mixed_v1
Pre-shuffled, pre-mixed training bundle for the Base tier (201.3M params)
of the contrastive forecasting model family.
Total rows: 106,100,279
Number of shards: 10,624 (9,985 in base_mixed_v1/, 639 in base_mixed_v1/overflow/)
Bundle size: ~397 GB
Per-source row counts
0 gift: 99,333,452 (93.6%)
1 wiki_hourly: 3,715,121 (3.5%)
2 wiki_daily: 1,990,244 (1.9%)
3 wiki_stl_residual: 530,731 (0.5%)
4 wiki_stl_seasonal: 371,512 (0.4%)
5 wiki_stl_trend: 159,219… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/contrastive-training-base-bundles.turkish_weakly_supervised_contrastive_learning_datasetcontrastive-training-small-bundles
small_mixed_v1
Pre-shuffled, pre-mixed training bundle for the Small tier (42.5M params)
of the contrastive forecasting model family.
Total rows: 27282687
Number of shards: 2736
Per-source row counts (source_id -> rows)
0 gift: 20250495 (74.2%)
1 wiki_hourly: 3715121 (13.6%)
2 wiki_daily: 1990244 (7.3%)
3 wiki_stl_residual: 530731 (1.9%)
4 wiki_stl_seasonal: 371512 (1.4%)
5 wiki_stl_trend: 159219 (0.6%)
6 synthetic: 265365 (1.0%)
Schema
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/contrastive-training-small-bundles.medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastive-with-choices-subsampled_trakmedmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastivemedmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastive-with-choicescontrastive-belief-updates
Contrastive SDF training corpora
This dataset is from Apollo Research and accompanies the paper Measuring Reward-Seeking via Contrastive
Belief Updates. For more, see rewardseeking.ai.
This dataset contains the 30 synthetic-document corpora used across the completed experiments for the paper:
24 coding-style corpora and 6 honesty-versus-task-completion corpora.
Important: entirely synthetic, model-generated content
Every document in this dataset is synthetic and… See the full description on the dataset page: https://huggingface.co/datasets/apollo-research/contrastive-belief-updates.dalle-3-contrastive-captions-updated
Dataset Card for "dalle-3-contrastive-captions-updated"
More Information needed
human_ai_story_contrastive_v4
Human–AI Story Contrastive Six-Source Dataset
Dataset summary
This dataset contains 1,443 matched short-story prompt groups with human and machine-written realizations drawn from six source populations.
The dataset is organized one row per group_id rather than one row per text. Each row preserves the common writing prompt, the human reference story, the available earlier machine generations, the starting-policy generation, and two GPT-5.6 Sol fields added for… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v4.mnli-mock-contrastive-axes-ii
Dataset Card for "mnli-mock-contrastive-axes-ii"
More Information needed
human_ai_story_contrastive_v2Collapsed and Augmented with llama3.3sft on-policy/policy adjacent responses
Prepare for some AI generated descriptions.
Dataset Construction and Statistics
Purpose: A contrastive creative-writing dataset for studying distributional stylistic differences between human-written and model-generated text.
Total size: 5,733 texts grouped across 1,443 unique prompts / human responses.
Source breakdown:
1,443 human responses
1,421 GPT-3.5 responses
1,426 Claude Opus responses
1,443… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v2.mnli-mock-contrastive-axes
Dataset Card for "mnli-mock-contrastive-axes"
More Information needed
doc-context-contrastive-triplets
Document Context Contrastive Triplets
This dataset provides training and validation examples designed for contrastive learning and representation tasks, specifically focusing on distinguishing contextually related text spans from unrelated spans.
Dataset Structure
Each example in the dataset contains the following fields:
Field
Type
Description
anchor
string
Random text span sampled from a source document.
positive
string
Random text span sampled… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/doc-context-contrastive-triplets.dalle-3-contrastive-captions
Dataset Card for "dalle-3-contrastive-captions"
More Information needed
Provenancenpedit-contrastive-edit-1.5M
NP-Edit Contrastive Instruction-Editing Dataset (1.5M)
Unpaired instruction-based image-editing data with contrastive source/target captions, assembled
for training few-step distilled image editors (DMD/VSD on SD3.5). Each example is a source image + an
edit instruction + a minimal-contrast (source caption, target caption) pair + a grounding noun for the
edited region. Edited/target images are intentionally not required (unpaired training).
Sources
Merged from two… See the full description on the dataset page: https://huggingface.co/datasets/quandao10/npedit-contrastive-edit-1.5M.BidirLM-Omni-Contrastive
BidirLM-Omni-Contrastive
This repository serves as the centralized documentation and integration hub for the omnimodal contrastive datasets used to train BidirLM-Omni and introduced in the paper BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs.
Rather than forcing a massive, monolithic download, this hub provides direct links to the modality-specific repositories and the exact Python code required to load, shuffle, and sample the data for… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/BidirLM-Omni-Contrastive.translation-contrastive-tripletsNemotron-Personas-Korea-Embedding-ContrastiveLearninglink_prediction_contrastivelibrispeech_contrastive
⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models.
Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive
📜 Citation
If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/librispeech_contrastive.
