CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /sst5 Stanford Sentiment Treebank - Fine-Grained Stanford Sentiment Treebank with 5 labels: very positive, positive, neutral, negative, very negative Splits are from: https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data Training data is on sentence level, not on phrase level! text10K<n<100K21 likes21k downloads5y agoHugging Face02SetFit /emotion** Attention: There appears an overlap in train / test. I trained a model on the train set and achieved 100% acc on test set. With the original emotion dataset this is not the case (92.4% acc)** text10K<n<100K30 likes13k downloads4y agoHugging Face03SetFit /20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs: The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date. We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.text10K<n<100K21 likes9.4k downloads5y agoHugging Face04SetFit /sst2 Stanford Sentiment Treebank - Binary Stanford Sentiment Treebank with 2 labels: negative, positive Splits are from: https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data Training data is on sentence level, not on phrase level! text1K<n<10K11 likes6.5k downloads5y agoHugging Face05SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes5.5k downloads5y agoHugging Face06SetFit /ag_newstext100K<n<1M10 likes5.1k downloads5y agoHugging Face07SetFit /rte Glue RTE This dataset is a port of the official rte dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K2 likes4.8k downloads5y agoHugging Face08SetFit /TREC-QC TREC Question Classification Question classification in coarse and fine-grained categories. Source: Experimental Data for Question Classification Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002. tabular1K<n<10K0 likes3k downloads5y agoHugging Face09SetFit /bbc-news BBC News Topic Dataset Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech. Original source for this dataset: Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning (ICML’06)… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/bbc-news.texttext-classification1K<n<10K24 likes2.7k downloads2y agoHugging Face10SetFit /CR Customer Reviews This dataset is a port of the official CR dataset from this paper. There is no validation split. text1K<n<10K2 likes2.6k downloads4y agoHugging Face11SetFit /amazon_counterfactual_en Amazon Counterfactual Statements This dataset is the en-ext split from SetFit/amazon_counterfactual. As the original test set is rather small (1333 examples), a different split was created with 50-50 for training & testing. The dataset is described in amazon-multilingual-counterfactual-dataset / Paper It contains statements from Amazon reviews about events that did not or cannot take place. text10K<n<100K0 likes2.5k downloads5y agoHugging Face12SetFit /amazon_reviews_multi_entext100K<n<1M7 likes2.4k downloads4y agoHugging Face13SetFit /subj Subjective vs Objective This is the SUBJ dataset as used in SentEval. It contains sentences with an annotation if they sentence describes something subjective about a movie or something objective text10K<n<100K8 likes2.3k downloads5y agoHugging Face14SetFit /SentEval-CR SentEval Customer Reviews This dataset is a port of the official SentEval CR dataset from this paper. The test split was created from the by randomly sampling 20% of the original data and the train split is the remaining 80%. there are no official train/test splits of CR. There is no validation split. This was used in the STraTA paper. text1K<n<10K3 likes2.3k downloads4y agoHugging Face15SetFit /ade_corpus_v2_classification ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data. This is a dataset for classification if a sentence is ADE-related (True) or not (False). Train size: 17,637 Test size: 5,879 Source dataset Paper text10K<n<100K6 likes1.8k downloads4y agoHugging Face16SetFit /mnli Glue MNLI This dataset is a port of the official mnli dataset on the Hub. It contains the matched version. Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M8 likes1.7k downloads5y agoHugging Face17SetFit /tweet_eval_stance_abortiontextn<1K0 likes1.7k downloads4y agoHugging Face18SetFit /qqp Glue QQP This dataset is a port of the official qqp dataset on the Hub. Note that the question1 and question2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M6 likes1.7k downloads5y agoHugging Face19SetFit /amazon_massive_intent_en-UStext10K<n<100K10 likes1.5k downloads4y agoHugging Face20SetFit /qnli Glue QNLI This dataset is a port of the official qnli dataset on the Hub. Note that the question and sentence columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M2 likes1.4k downloads5y agoHugging Face21SetFit /mrpc Glue MRPC This dataset is a port of the official mrpc dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K17 likes1.3k downloads5y agoHugging Face22argilla /banking_sentiment_setfit Dataset Card for "banking_sentiment_setfit" More Information needed textn<1K2 likes942 downloads4y agoHugging Face23SetFit /yelp_review_fulltext100K<n<1M1 likes921 downloads5y agoHugging Face24SetFit /imdbtext10K<n<100K3 likes808 downloads5y agoHugging Face25SetFit /go_emotions GoEmotions This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification. tabular10K<n<100K13 likes745 downloads4y agoHugging Face26SetFit /stsb Glue STS-B This dataset is a port of the official sts-b dataset on the Hub. This is not a classification task, so the label_text column is only included for consistency Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K1 likes689 downloads5y agoHugging Face27SetFit /xglue_nc#xglue nc This dataset is a port of the official ['xglue' dataset] (https://huggingface.co/datasets/xglue) on the Hub. It has just the news category classification section. It has been reduced to just 3 columns (plus text label) that are relevant to the SetFit task. Validation and test in English, Spanish, French, Russian, and German. tabular100K<n<1M0 likes656 downloads2y agoHugging Face28SetFit /hate_speech_offensive hate_speech_offensive This dataset is a version from hate_speech_offensive, splitted into train and test set. text10K<n<100K2 likes634 downloads5y agoHugging Face29SetFit /toxic_conversations Toxic Conversation This is a version of the Jigsaw Unintended Bias in Toxicity Classification dataset. It contains comments from the Civil Comments platform together with annotations if the comment is toxic or not. 10 annotators annotated each example and, as recommended in the task page, set a comment as toxic when target >= 0.5 The dataset is inbalanced, with only about 8% of the comments marked as toxic. text1M<n<10M16 likes604 downloads5y agoHugging Face30SetFit /amazon_reviews_multi_ja#amazon reviews multi japanese This dataset is a port of the official ['amazon_reviews_multi' dataset] (https://huggingface.co/datasets/amazon_reviews_multi) on the Hub. It has just the Japanese language version. It has been reduced to just 3 columns (and 4th "label_text") that are relevant to the SetFit task. text100K<n<1M7 likes553 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.