contrastive
contrastive-stubscontrastive-probing-macmmarco-contrastive
mMARCO-contrastive
The dataset is a modification of mMARCO focusing on French and English parts. The aim is to train a
bi-encoder model using all hard negatives from the database. Instead of having a query/positive/negative triplet, we pair all negatives with a query and a
positive. However, it's worth noting that there are many false negatives in the dataset. This isn't a big issue with a triplet view because false negatives
are much fewer in number, but it's more significant with… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/mmarco-contrastive.laion_audio_contrastive
⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models.
Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive
📜 Citation
If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/laion_audio_contrastive.BidirLM-Contrastive
BidirLM-Contrastive
The contrastive training dataset used to train BidirLM Embedding models. It contains 10,110,219 query-document pairs from 79 base datasets, split into 203 subdatasets by language or type (~13 GB), covering three sources: Nemotron, KaLM, and parallel/other data. This dataset is described in the paper: BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs.
If you use this dataset in your research or applications, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/BidirLM-Contrastive.halvest-contrastive
HALvest-Contrastive
Contrastive triplets Harvested from HAL
Citation
@misc{kulumba2026doesauthorshipsignalemerge,
title={Where Does Authorship Signal Emerge in Encoder-Based Language Models?},
author={Francis Kulumba and Guillaume Vimont and Laurent Romary and Florian Cafiero},
year={2026},
eprint={2605.19908},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2605.19908},
}… See the full description on the dataset page: https://huggingface.co/datasets/almanach/halvest-contrastive.
