CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AST-Revisited /contrastive-stubsimage0 likes1.1k downloads6mo agoHugging Face02Cyrile /mmarco-contrastive mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to train a bi-encoder model using all hard negatives from the database. Instead of having a query/positive/negative triplet, we pair all negatives with a query and a positive. However, it's worth noting that there are many false negatives in the dataset. This isn't a big issue with a triplet view because false negatives are much fewer in number, but it's more significant with… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/mmarco-contrastive.texttranslation100K<n<1M2 likes715 downloads2y agoHugging Face03ch-min /contrastive-probing-macimage10K<n<100K0 likes523 downloads5mo agoHugging Face04orionweller /contrastive-pretraining Contrastive Pretraining Per-language query/document pairs produced by the retrieval-common-crawl pipeline. Each config corresponds to a single language or source with identical LightOn-style schema. Config overview Configs available are: fw-edu, fw2-arb_Arab, fw2-ces_Latn, fw2-cmn_Hani, fw2-dan_Latn, fw2-deu_Latn, fw2-ell_Grek, fw2-fas_Arab, fw2-fra_Latn, fw2-hun_Latn, fw2-ind_Latn, fw2-ita_Latn, fw2-jpn_Jpan, fw2-nld_Latn, fw2-pol_Latn, fw2-por_Latn, fw2-rus_Cyrl… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/contrastive-pretraining.tabular100M<n<1B3 likes517 downloads5mo agoHugging Face05BidirLM /laion_audio_contrastive ⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models. Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive 📜 Citation If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/laion_audio_contrastive.text100K<n<1M0 likes479 downloads4mo agoHugging Face06BidirLM /BidirLM-Contrastive BidirLM-Contrastive The contrastive training dataset used to train BidirLM Embedding models. It contains 10,110,219 query-document pairs from 79 base datasets, split into 203 subdatasets by language or type (~13 GB), covering three sources: Nemotron, KaLM, and parallel/other data. This dataset is described in the paper: BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs. If you use this dataset in your research or applications, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/BidirLM-Contrastive.texttext-retrieval10M<n<100M3 likes454 downloads4mo agoHugging Face07almanach /halvest-contrastive HALvest-Contrastive Contrastive triplets Harvested from HAL Citation @misc{kulumba2026doesauthorshipsignalemerge, title={Where Does Authorship Signal Emerge in Encoder-Based Language Models?}, author={Francis Kulumba and Guillaume Vimont and Laurent Romary and Florian Cafiero}, year={2026}, eprint={2605.19908}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2605.19908}, }… See the full description on the dataset page: https://huggingface.co/datasets/almanach/halvest-contrastive.texttext-classification1M<n<10M1 likes378 downloads4mo agoHugging Face08BidirLM /mscoco_contrastive ⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models. Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive 📜 Citation If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/mscoco_contrastive.image100K<n<1M0 likes264 downloads4mo agoHugging Face09almanach /halvest-contrastive-oldtext1M<n<10M0 likes235 downloads1y agoHugging Face10jeremycochoy /contrastive-training-base-bundles base_mixed_v1 Pre-shuffled, pre-mixed training bundle for the Base tier (201.3M params) of the contrastive forecasting model family. Total rows: 106,100,279 Number of shards: 10,624 (9,985 in base_mixed_v1/, 639 in base_mixed_v1/overflow/) Bundle size: ~397 GB Per-source row counts 0 gift: 99,333,452 (93.6%) 1 wiki_hourly: 3,715,121 (3.5%) 2 wiki_daily: 1,990,244 (1.9%) 3 wiki_stl_residual: 530,731 (0.5%) 4 wiki_stl_seasonal: 371,512 (0.4%) 5 wiki_stl_trend: 159,219… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/contrastive-training-base-bundles.text100M<n<1B0 likes224 downloads5mo agoHugging Face11selmanbaysan /turkish_weakly_supervised_contrastive_learning_datasettext10M<n<100M0 likes193 downloads1y agoHugging Face12jeremycochoy /contrastive-training-small-bundles small_mixed_v1 Pre-shuffled, pre-mixed training bundle for the Small tier (42.5M params) of the contrastive forecasting model family. Total rows: 27282687 Number of shards: 2736 Per-source row counts (source_id -> rows) 0 gift: 20250495 (74.2%) 1 wiki_hourly: 3715121 (13.6%) 2 wiki_daily: 1990244 (7.3%) 3 wiki_stl_residual: 530731 (1.9%) 4 wiki_stl_seasonal: 371512 (1.4%) 5 wiki_stl_trend: 159219 (0.6%) 6 synthetic: 265365 (1.0%) Schema Column Type… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/contrastive-training-small-bundles.text10M<n<100M0 likes152 downloads5mo agoHugging Face13zekeZZ /medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastive-with-choices-subsampled_traktextn<1K0 likes137 downloads2y agoHugging Face14zekeZZ /medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastivetext1K<n<10K0 likes118 downloads2y agoHugging Face15zekeZZ /medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastive-with-choicestext1K<n<10K0 likes109 downloads2y agoHugging Face16apollo-research /contrastive-belief-updates Contrastive SDF training corpora This dataset is from Apollo Research and accompanies the paper Measuring Reward-Seeking via Contrastive Belief Updates. For more, see rewardseeking.ai. This dataset contains the 30 synthetic-document corpora used across the completed experiments for the paper: 24 coding-style corpora and 6 honesty-versus-task-completion corpora. Important: entirely synthetic, model-generated content Every document in this dataset is synthetic and… See the full description on the dataset page: https://huggingface.co/datasets/apollo-research/contrastive-belief-updates.text100K<n<1M4 likes99 downloads2mo agoHugging Face17facet /dalle-3-contrastive-captions-updated Dataset Card for "dalle-3-contrastive-captions-updated" More Information needed image1K<n<10K1 likes88 downloads3y agoHugging Face18schonsense /human_ai_story_contrastive_v4 Human–AI Story Contrastive Six-Source Dataset Dataset summary This dataset contains 1,443 matched short-story prompt groups with human and machine-written realizations drawn from six source populations. The dataset is organized one row per group_id rather than one row per text. Each row preserves the common writing prompt, the human reference story, the available earlier machine generations, the starting-policy generation, and two GPT-5.6 Sol fields added for… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v4.text1K<n<10K0 likes80 downloads4d agoHugging Face19iamroot /mnli-mock-contrastive-axes-ii Dataset Card for "mnli-mock-contrastive-axes-ii" More Information needed text100K<n<1M0 likes77 downloads3y agoHugging Face20schonsense /human_ai_story_contrastive_v2Collapsed and Augmented with llama3.3sft on-policy/policy adjacent responses Prepare for some AI generated descriptions. Dataset Construction and Statistics Purpose: A contrastive creative-writing dataset for studying distributional stylistic differences between human-written and model-generated text. Total size: 5,733 texts grouped across 1,443 unique prompts / human responses. Source breakdown: 1,443 human responses 1,421 GPT-3.5 responses 1,426 Claude Opus responses 1,443… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v2.text1K<n<10K0 likes73 downloads15d agoHugging Face21iamroot /mnli-mock-contrastive-axes Dataset Card for "mnli-mock-contrastive-axes" More Information needed text100K<n<1M0 likes71 downloads3y agoHugging Face22agentlans /doc-context-contrastive-triplets Document Context Contrastive Triplets This dataset provides training and validation examples designed for contrastive learning and representation tasks, specifically focusing on distinguishing contextually related text spans from unrelated spans. Dataset Structure Each example in the dataset contains the following fields: Field Type Description anchor string Random text span sampled from a source document. positive string Random text span sampled… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/doc-context-contrastive-triplets.textfeature-extraction100K<n<1M0 likes68 downloads1mo agoHugging Face23facet /dalle-3-contrastive-captions Dataset Card for "dalle-3-contrastive-captions" More Information needed image1K<n<10K0 likes66 downloads3y agoHugging Face24Contrastive-Prov /Provenance0 likes63 downloads2y agoHugging Face25quandao10 /npedit-contrastive-edit-1.5M NP-Edit Contrastive Instruction-Editing Dataset (1.5M) Unpaired instruction-based image-editing data with contrastive source/target captions, assembled for training few-step distilled image editors (DMD/VSD on SD3.5). Each example is a source image + an edit instruction + a minimal-contrast (source caption, target caption) pair + a grounding noun for the edited region. Edited/target images are intentionally not required (unpaired training). Sources Merged from two… See the full description on the dataset page: https://huggingface.co/datasets/quandao10/npedit-contrastive-edit-1.5M.image-to-image1M<n<10M0 likes61 downloads11d agoHugging Face26BidirLM /BidirLM-Omni-Contrastive BidirLM-Omni-Contrastive This repository serves as the centralized documentation and integration hub for the omnimodal contrastive datasets used to train BidirLM-Omni and introduced in the paper BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs. Rather than forcing a massive, monolithic download, this hub provides direct links to the modality-specific repositories and the exact Python code required to load, shuffle, and sample the data for… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/BidirLM-Omni-Contrastive.feature-extraction1M<n<10M5 likes59 downloads4mo agoHugging Face27agentlans /translation-contrastive-tripletstext100K<n<1M0 likes59 downloads14d agoHugging Face28hyunseop /Nemotron-Personas-Korea-Embedding-ContrastiveLearning1 likes56 downloads5mo agoHugging Face29ngarneau /link_prediction_contrastivetext100K<n<1M5 likes54 downloads3y agoHugging Face30BidirLM /librispeech_contrastive ⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models. Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive 📜 Citation If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/librispeech_contrastive.audio100K<n<1M0 likes52 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.