datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham.
rte
Glue RTE
This dataset is a port of the official rte dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
TREC-QC
TREC Question Classification
Question classification in coarse and fine-grained categories.
Source:
Experimental Data for Question Classification
Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002.
mnli
Glue MNLI
This dataset is a port of the official mnli dataset on the Hub.
It contains the matched version.
Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
qqp
Glue QQP
This dataset is a port of the official qqp dataset on the Hub.
Note that the question1 and question2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
qnli
Glue QNLI
This dataset is a port of the official qnli dataset on the Hub.
Note that the question and sentence columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
mrpc
Glue MRPC
This dataset is a port of the official mrpc dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
stsb
Glue STS-B
This dataset is a port of the official sts-b dataset on the Hub.
This is not a classification task, so the label_text column is only included for consistency
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
xglue_nc#xglue nc
This dataset is a port of the official ['xglue' dataset] (https://huggingface.co/datasets/xglue) on the Hub. It has just the news category classification section. It has been reduced to just 3 columns (plus text label) that are relevant to the SetFit task. Validation and test in English, Spanish, French, Russian, and German.
hate_speech18wsc_fixed
Glue WSC Fixed
This dataset is a port of the official wsc.fixed dataset on the Hub.
Also, the test split is not labeled; the label column values are always -1.
wnli
Glue WNLI
This dataset is a port of the official wnli dataset on the Hub.
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
wsc
Glue WSC
This dataset is a port of the official wsc dataset on the Hub.
Also, the test split is not labeled; the label column values are always -1.
mnli_mm
Glue MNLI
This dataset is a port of the official mnli dataset on the Hub.
It contains the mismatched version.
Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
setfit-absa-tesla-tweetsMAMS_ACSA_SETFITABSAsetfit-proj8-multilabel_2MAMS_ATSA_SETFITABSAsetfit-proj8-multilabel_2_validationstsb-setfit-copySetFit_mnli-copy
