CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cyrile /mmarco-contrastive mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to train a bi-encoder model using all hard negatives from the database. Instead of having a query/positive/negative triplet, we pair all negatives with a query and a positive. However, it's worth noting that there are many false negatives in the dataset. This isn't a big issue with a triplet view because false negatives are much fewer in number, but it's more significant with… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/mmarco-contrastive.texttranslation100K<n<1M2 likes660 downloads2y agoHugging Face02BidirLM /laion_audio_contrastive ⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models. Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive 📜 Citation If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/laion_audio_contrastive.text100K<n<1M0 likes479 downloads4mo agoHugging Face03orionweller /contrastive-pretraining Contrastive Pretraining Per-language query/document pairs produced by the retrieval-common-crawl pipeline. Each config corresponds to a single language or source with identical LightOn-style schema. Config overview Configs available are: fw-edu, fw2-arb_Arab, fw2-ces_Latn, fw2-cmn_Hani, fw2-dan_Latn, fw2-deu_Latn, fw2-ell_Grek, fw2-fas_Arab, fw2-fra_Latn, fw2-hun_Latn, fw2-ind_Latn, fw2-ita_Latn, fw2-jpn_Jpan, fw2-nld_Latn, fw2-pol_Latn, fw2-por_Latn, fw2-rus_Cyrl… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/contrastive-pretraining.tabular100M<n<1B3 likes463 downloads5mo agoHugging Face04BidirLM /BidirLM-Contrastive BidirLM-Contrastive The contrastive training dataset used to train BidirLM Embedding models. It contains 10,110,219 query-document pairs from 79 base datasets, split into 203 subdatasets by language or type (~13 GB), covering three sources: Nemotron, KaLM, and parallel/other data. This dataset is described in the paper: BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs. If you use this dataset in your research or applications, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/BidirLM-Contrastive.texttext-retrieval10M<n<100M3 likes451 downloads4mo agoHugging Face05almanach /halvest-contrastive HALvest-Contrastive Contrastive triplets Harvested from HAL Citation @misc{kulumba2026doesauthorshipsignalemerge, title={Where Does Authorship Signal Emerge in Encoder-Based Language Models?}, author={Francis Kulumba and Guillaume Vimont and Laurent Romary and Florian Cafiero}, year={2026}, eprint={2605.19908}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2605.19908}, }… See the full description on the dataset page: https://huggingface.co/datasets/almanach/halvest-contrastive.texttext-classification1M<n<10M1 likes376 downloads4mo agoHugging Face06almanach /halvest-contrastive-oldtext1M<n<10M0 likes249 downloads1y agoHugging Face07jeremycochoy /contrastive-training-base-bundles base_mixed_v1 Pre-shuffled, pre-mixed training bundle for the Base tier (201.3M params) of the contrastive forecasting model family. Total rows: 106,100,279 Number of shards: 10,624 (9,985 in base_mixed_v1/, 639 in base_mixed_v1/overflow/) Bundle size: ~397 GB Per-source row counts 0 gift: 99,333,452 (93.6%) 1 wiki_hourly: 3,715,121 (3.5%) 2 wiki_daily: 1,990,244 (1.9%) 3 wiki_stl_residual: 530,731 (0.5%) 4 wiki_stl_seasonal: 371,512 (0.4%) 5 wiki_stl_trend: 159,219… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/contrastive-training-base-bundles.text100M<n<1B0 likes224 downloads5mo agoHugging Face08BidirLM /mscoco_contrastive ⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models. Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive 📜 Citation If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/mscoco_contrastive.image100K<n<1M0 likes204 downloads4mo agoHugging Face09selmanbaysan /turkish_weakly_supervised_contrastive_learning_datasettext10M<n<100M0 likes198 downloads1y agoHugging Face10jeremycochoy /contrastive-training-small-bundles small_mixed_v1 Pre-shuffled, pre-mixed training bundle for the Small tier (42.5M params) of the contrastive forecasting model family. Total rows: 27282687 Number of shards: 2736 Per-source row counts (source_id -> rows) 0 gift: 20250495 (74.2%) 1 wiki_hourly: 3715121 (13.6%) 2 wiki_daily: 1990244 (7.3%) 3 wiki_stl_residual: 530731 (1.9%) 4 wiki_stl_seasonal: 371512 (1.4%) 5 wiki_stl_trend: 159219 (0.6%) 6 synthetic: 265365 (1.0%) Schema Column Type… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/contrastive-training-small-bundles.text10M<n<100M0 likes152 downloads5mo agoHugging Face11zekeZZ /medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastive-with-choices-subsampled_traktextn<1K0 likes138 downloads2y agoHugging Face12zekeZZ /medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastivetext1K<n<10K0 likes119 downloads2y agoHugging Face13zekeZZ /medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastive-with-choicestext1K<n<10K0 likes109 downloads2y agoHugging Face14apollo-research /contrastive-belief-updates Contrastive SDF training corpora This dataset is from Apollo Research and accompanies the paper Measuring Reward-Seeking via Contrastive Belief Updates. For more, see rewardseeking.ai. This dataset contains the 30 synthetic-document corpora used across the completed experiments for the paper: 24 coding-style corpora and 6 honesty-versus-task-completion corpora. Important: entirely synthetic, model-generated content Every document in this dataset is synthetic and… See the full description on the dataset page: https://huggingface.co/datasets/apollo-research/contrastive-belief-updates.text100K<n<1M4 likes99 downloads2mo agoHugging Face15facet /dalle-3-contrastive-captions-updated Dataset Card for "dalle-3-contrastive-captions-updated" More Information needed image1K<n<10K1 likes88 downloads3y agoHugging Face16schonsense /human_ai_story_contrastive_v4 Human–AI Story Contrastive Six-Source Dataset Dataset summary This dataset contains 1,443 matched short-story prompt groups with human and machine-written realizations drawn from six source populations. The dataset is organized one row per group_id rather than one row per text. Each row preserves the common writing prompt, the human reference story, the available earlier machine generations, the starting-policy generation, and two GPT-5.6 Sol fields added for… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v4.text1K<n<10K0 likes82 downloads6d agoHugging Face17iamroot /mnli-mock-contrastive-axes-ii Dataset Card for "mnli-mock-contrastive-axes-ii" More Information needed text100K<n<1M0 likes77 downloads3y agoHugging Face18schonsense /human_ai_story_contrastive_v2Collapsed and Augmented with llama3.3sft on-policy/policy adjacent responses Prepare for some AI generated descriptions. Dataset Construction and Statistics Purpose: A contrastive creative-writing dataset for studying distributional stylistic differences between human-written and model-generated text. Total size: 5,733 texts grouped across 1,443 unique prompts / human responses. Source breakdown: 1,443 human responses 1,421 GPT-3.5 responses 1,426 Claude Opus responses 1,443… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v2.text1K<n<10K0 likes75 downloads16d agoHugging Face19iamroot /mnli-mock-contrastive-axes Dataset Card for "mnli-mock-contrastive-axes" More Information needed text100K<n<1M0 likes71 downloads3y agoHugging Face20facet /dalle-3-contrastive-captions Dataset Card for "dalle-3-contrastive-captions" More Information needed image1K<n<10K0 likes66 downloads3y agoHugging Face21agentlans /translation-contrastive-tripletstext100K<n<1M0 likes59 downloads15d agoHugging Face22agentlans /doc-context-contrastive-triplets Document Context Contrastive Triplets This dataset provides training and validation examples designed for contrastive learning and representation tasks, specifically focusing on distinguishing contextually related text spans from unrelated spans. Dataset Structure Each example in the dataset contains the following fields: Field Type Description anchor string Random text span sampled from a source document. positive string Random text span sampled… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/doc-context-contrastive-triplets.textfeature-extraction100K<n<1M0 likes55 downloads1mo agoHugging Face23ngarneau /link_prediction_contrastivetext100K<n<1M6 likes54 downloads3y agoHugging Face24ozertuu /turkish-food-contrastive-triplets 🥘 Turkish & Multi-Lingual Food Contrastive Triplets Dataset This dataset contains 30,001 culinary triplets (Anchor, Positive, Hard Negative) designed for training contrastive dense retrieval models, embedding projection heads, and fine-grained food similarity metrics. Used for training the ozertuu/bge-m3-food-projection-head. 📊 Dataset Structure Each record consists of: anchor (string): The base query or canonical food name (e.g., "Kuru Fasulye"). positive… See the full description on the dataset page: https://huggingface.co/datasets/ozertuu/turkish-food-contrastive-triplets.text10K<n<100K0 likes53 downloads23d agoHugging Face25BidirLM /librispeech_contrastive ⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models. Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive 📜 Citation If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/librispeech_contrastive.audio100K<n<1M0 likes52 downloads4mo agoHugging Face26Lexsi /circuitkit-capitals-contrastive CircuitKIT capitals — contrastive pairs Twelve capital-city facts, each with an explicit counterfactual pair, for circuit discovery with CircuitKIT. column meaning question clean prompt, e.g. The capital of France is answer clean answer, e.g. Paris corrupted_question counterfactual prompt of the same shape, e.g. The capital of Germany is corrupted_answer its answer, e.g. Berlin Attribution-patching methods (EAP, EAP-IG, …) score a component by how much it… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/circuitkit-capitals-contrastive.texttext-generationn<1K0 likes52 downloads4d agoHugging Face27annnli /C4-contrastive-watermark Dataset Card for "C4-contrastive-watermark" More Information needed tabular1K<n<10K0 likes51 downloads1y agoHugging Face28mjbommar /opengloss-v1.3-contrastive-examples See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Contrastive Examples v1.3 Dataset Summary OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations designed for contrastive learning and… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-contrastive-examples.tabulartext-generation100K<n<1M0 likes51 downloads16d agoHugging Face29Samll /Llama-3.1-8B-Instruct_Taur_CoT_contrastive_relevancestext10K<n<100K0 likes50 downloads2y agoHugging Face30reasoning-degeneration-dev /gepa-rlm-exp-contrastive-20260219-191545 gepa-rlm-exp-contrastive-20260219-191545 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: contrastive | Last updated: 2026-02-19 21:46 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_rlm_k20 rlm 20 contrastive 48.89% 33.33% 826,936 $0.0000 6170s Learning Curves Experiment Config { "script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-contrastive-20260219-191545.tabularn<1K0 likes48 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.