CoolFace
20 results

simcse

princeton-nlp /datasets-for-simcsetext1M<n<10M8 likes1k downloads5y agoHugging Faceclosji /bookcorpus_filtered_len_17_simcsetext10M<n<100M0 likes332 downloads4y agoHugging Faceclosji /wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largetext1M<n<10M0 likes186 downloads4y agoHugging Facesentence-transformers /nli-for-simcse Dataset Card for NLI for SimCSE This is a reformatting of the NLI for SimCSE Dataset used to train the BGE-M3 model. See the full BGE-M3 dataset in Shitao/bge-m3-data. Despite being labeled as Natural Language Inference (NLI), this dataset can be used for training/finetuning an embedding model for semantic textual similarity. Dataset Subsets triplet subset Columns: "anchor", "positive", "negative" Column types: str, str, str Examples:{ 'anchor': 'One… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/nli-for-simcse.textfeature-extraction1M<n<10M3 likes153 downloads2y agoHugging Faceclosji /cc12m_princeton-nlp_sup-simcse-roberta-largeimage10M<n<100M0 likes110 downloads4y agoHugging Facedkoterwa /kor_nli_simcse Korean Natural Language Inference (KorNLI) for SimCSE Dataset For a better dataset description, please visit this GitHub repository prepared by the authors of the article: LINK This dataset was prepared by converting KorNLI dataset. I took every unique premise of the dataset and searched for its entailment (positive example) and contradiction (negative example). These changes have been made in order to apply SimCSE method. I additionaly share the code, which I used to convert the… See the full description on the dataset page: https://huggingface.co/datasets/dkoterwa/kor_nli_simcse.text100K<n<1M1 likes103 downloads3y agoHugging Face