CoolFace
20 results

contrastive

AST-Revisited /contrastive-stubsimage0 likes1.1k downloads6mo agoHugging FaceCyrile /mmarco-contrastive mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to train a bi-encoder model using all hard negatives from the database. Instead of having a query/positive/negative triplet, we pair all negatives with a query and a positive. However, it's worth noting that there are many false negatives in the dataset. This isn't a big issue with a triplet view because false negatives are much fewer in number, but it's more significant with… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/mmarco-contrastive.texttranslation100K<n<1M2 likes715 downloads2y agoHugging Facech-min /contrastive-probing-macimage10K<n<100K0 likes523 downloads5mo agoHugging Faceorionweller /contrastive-pretraining Contrastive Pretraining Per-language query/document pairs produced by the retrieval-common-crawl pipeline. Each config corresponds to a single language or source with identical LightOn-style schema. Config overview Configs available are: fw-edu, fw2-arb_Arab, fw2-ces_Latn, fw2-cmn_Hani, fw2-dan_Latn, fw2-deu_Latn, fw2-ell_Grek, fw2-fas_Arab, fw2-fra_Latn, fw2-hun_Latn, fw2-ind_Latn, fw2-ita_Latn, fw2-jpn_Jpan, fw2-nld_Latn, fw2-pol_Latn, fw2-por_Latn, fw2-rus_Cyrl… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/contrastive-pretraining.tabular100M<n<1B3 likes517 downloads5mo agoHugging FaceBidirLM /laion_audio_contrastive ⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models. Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive 📜 Citation If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/laion_audio_contrastive.text100K<n<1M0 likes479 downloads4mo agoHugging FaceBidirLM /BidirLM-Contrastive BidirLM-Contrastive The contrastive training dataset used to train BidirLM Embedding models. It contains 10,110,219 query-document pairs from 79 base datasets, split into 203 subdatasets by language or type (~13 GB), covering three sources: Nemotron, KaLM, and parallel/other data. This dataset is described in the paper: BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs. If you use this dataset in your research or applications, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/BidirLM-Contrastive.texttext-retrieval10M<n<100M3 likes454 downloads4mo agoHugging Face