CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01closji /bookcorpus_filtered_len_17_simcsetext10M<n<100M0 likes382 downloads4y agoHugging Face02closji /wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largetext1M<n<10M0 likes211 downloads4y agoHugging Face03sentence-transformers /nli-for-simcse Dataset Card for NLI for SimCSE This is a reformatting of the NLI for SimCSE Dataset used to train the BGE-M3 model. See the full BGE-M3 dataset in Shitao/bge-m3-data. Despite being labeled as Natural Language Inference (NLI), this dataset can be used for training/finetuning an embedding model for semantic textual similarity. Dataset Subsets triplet subset Columns: "anchor", "positive", "negative" Column types: str, str, str Examples:{ 'anchor': 'One… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/nli-for-simcse.textfeature-extraction1M<n<10M3 likes149 downloads2y agoHugging Face04closji /cc12m_princeton-nlp_sup-simcse-roberta-largeimage10M<n<100M0 likes126 downloads4y agoHugging Face05dkoterwa /kor_nli_simcse Korean Natural Language Inference (KorNLI) for SimCSE Dataset For a better dataset description, please visit this GitHub repository prepared by the authors of the article: LINK This dataset was prepared by converting KorNLI dataset. I took every unique premise of the dataset and searched for its entailment (positive example) and contradiction (negative example). These changes have been made in order to apply SimCSE method. I additionaly share the code, which I used to convert the… See the full description on the dataset page: https://huggingface.co/datasets/dkoterwa/kor_nli_simcse.text100K<n<1M1 likes102 downloads3y agoHugging Face06llm-book /jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3 Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3" More Information needed tabular1M<n<10M1 likes85 downloads3y agoHugging Face07sentence-transformers /wiki1m-for-simcse Dataset Card for Wiki1m for SimCSE This is a reupload of the wiki1m_for_simcse.txt file from princeton-nlp/datasets-for-simcse, which can no longer be downloaded with recent datasets versions. Columns: "text" Column types: str Examples:{'text': 'YMCA in South Australia'} Collection strategy: Downloading the princeton-nlp/datasets-for-simcse dataset with datasets==2.21.0 and reuploading it to make the format compatible with datasets. Deduplicated: No textfeature-extraction1M<n<10M1 likes69 downloads8mo agoHugging Face08closji /flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similarity_copytabular10M<n<100M0 likes55 downloads4y agoHugging Face09closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_10__target_tranch_10__from_120text100K<n<1M0 likes44 downloads4y agoHugging Face10closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_19__target_tranch_26__from_120text100K<n<1M0 likes37 downloads4y agoHugging Face11closji /flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similaritytabular10M<n<100M1 likes36 downloads4y agoHugging Face12closji /mscoco_2014_captions_princeton-nlp_sup-simcse-roberta-largetext100K<n<1M0 likes35 downloads4y agoHugging Face13closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_14__target_tranch_9__from_120text100K<n<1M0 likes34 downloads4y agoHugging Face14closji /multipit_crowd_all_count_simcse_retrieval_pairstabular1M<n<10M0 likes33 downloads4y agoHugging Face15closji /flickr30k_captions_simCSEtext100K<n<1M0 likes31 downloads4y agoHugging Face16closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_13__target_tranch_19__from_120text100K<n<1M0 likes30 downloads4y agoHugging Face17closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_13__target_tranch_16__from_120text100K<n<1M0 likes30 downloads4y agoHugging Face18closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_13__target_tranch_9__from_120text100K<n<1M0 likes29 downloads4y agoHugging Face19phnyxlab /klue-nli-simcse KLUENLI for SimCSE Dataset For a better dataset description, please visit: LINK This dataset was prepared by converting KLUENLI dataset to use it for contrastive training (SimCSE). The code used to prepare the data is given below: import pandas as pd from datasets import load_dataset, concatenate_datasets, Dataset from torch.utils.data import random_split class PrepTriplets: @staticmethod def make_dataset(): train_dataset = load_dataset("klue", "nli"… See the full description on the dataset page: https://huggingface.co/datasets/phnyxlab/klue-nli-simcse.text1K<n<10K1 likes29 downloads2y agoHugging Face20closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_19__target_tranch_27__from_120text100K<n<1M0 likes28 downloads4y agoHugging Face21yehzw /simcsetext1M<n<10M0 likes28 downloads10mo agoHugging Face22closji /multipit_crowd_all_count_SimCSEtext100K<n<1M0 likes23 downloads4y agoHugging Face23closji /multipit_auto_SimCSEtext100K<n<1M0 likes23 downloads4y agoHugging Face24closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_19__target_tranch_28__from_120text100K<n<1M0 likes23 downloads4y agoHugging Face25SeanLee97 /nli_for_simcsetext100K<n<1M0 likes23 downloads2y agoHugging Face26closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_10__target_tranch_28__from_120text100K<n<1M0 likes21 downloads4y agoHugging Face27closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_10__target_tranch_2__from_120text100K<n<1M0 likes21 downloads4y agoHugging Face28closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_11__target_tranch_3__from_120text100K<n<1M0 likes21 downloads4y agoHugging Face29closji /flickr30k_princeton-nlp_sup-simcse-roberta-largetext100K<n<1M0 likes19 downloads4y agoHugging Face30closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_12__target_tranch_9__from_120text100K<n<1M0 likes19 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.