CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /sst5 Stanford Sentiment Treebank - Fine-Grained Stanford Sentiment Treebank with 5 labels: very positive, positive, neutral, negative, very negative Splits are from: https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data Training data is on sentence level, not on phrase level! text10K<n<100K21 likes24k downloads5y agoHugging Face02SetFit /emotion** Attention: There appears an overlap in train / test. I trained a model on the train set and achieved 100% acc on test set. With the original emotion dataset this is not the case (92.4% acc)** text10K<n<100K30 likes14k downloads4y agoHugging Face03SetFit /20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs: The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date. We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.text10K<n<100K21 likes9.4k downloads5y agoHugging Face04SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes6.3k downloads5y agoHugging Face05SetFit /sst2 Stanford Sentiment Treebank - Binary Stanford Sentiment Treebank with 2 labels: negative, positive Splits are from: https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data Training data is on sentence level, not on phrase level! text1K<n<10K11 likes6.3k downloads5y agoHugging Face06SetFit /ag_newstext100K<n<1M10 likes5.1k downloads5y agoHugging Face07SetFit /rte Glue RTE This dataset is a port of the official rte dataset on the Hub. Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular1K<n<10K2 likes4.8k downloads5y agoHugging Face08SetFit /TREC-QC TREC Question Classification Question classification in coarse and fine-grained categories. Source: Experimental Data for Question Classification Xin Li, Dan Roth, Learning Question Classifiers. COLING'02, Aug., 2002. tabular1K<n<10K0 likes3.3k downloads5y agoHugging Face09SetFit /bbc-news BBC News Topic Dataset Dataset on BBC News Topic Classification consisting of 2,225 articles published on the BBC News website corresponding during 2004-2005. Each article is labeled under one of 5 categories: business, entertainment, politics, sport or tech. Original source for this dataset: Derek Greene, Pádraig Cunningham, “Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering,” in Proc. 23rd International Conference on Machine learning (ICML’06)… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/bbc-news.texttext-classification1K<n<10K24 likes2.9k downloads2y agoHugging Face10SetFit /amazon_reviews_multi_entext100K<n<1M7 likes2.8k downloads4y agoHugging Face11SetFit /amazon_counterfactual_en Amazon Counterfactual Statements This dataset is the en-ext split from SetFit/amazon_counterfactual. As the original test set is rather small (1333 examples), a different split was created with 50-50 for training & testing. The dataset is described in amazon-multilingual-counterfactual-dataset / Paper It contains statements from Amazon reviews about events that did not or cannot take place. text10K<n<100K0 likes2.6k downloads5y agoHugging Face12SetFit /CR Customer Reviews This dataset is a port of the official CR dataset from this paper. There is no validation split. text1K<n<10K2 likes2.6k downloads4y agoHugging Face13SetFit /SentEval-CR SentEval Customer Reviews This dataset is a port of the official SentEval CR dataset from this paper. The test split was created from the by randomly sampling 20% of the original data and the train split is the remaining 80%. there are no official train/test splits of CR. There is no validation split. This was used in the STraTA paper. text1K<n<10K3 likes2.4k downloads4y agoHugging Face14SetFit /subj Subjective vs Objective This is the SUBJ dataset as used in SentEval. It contains sentences with an annotation if they sentence describes something subjective about a movie or something objective text10K<n<100K8 likes2.3k downloads5y agoHugging Face15lerobot-raw /robo_set_rawn<1K0 likes2.2k downloads2y agoHugging Face16SetFit /ade_corpus_v2_classification ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data. This is a dataset for classification if a sentence is ADE-related (True) or not (False). Train size: 17,637 Test size: 5,879 Source dataset Paper text10K<n<100K6 likes1.9k downloads4y agoHugging Face17nc1708 /sn38-set-ctextn<1K0 likes1.7k downloads11d agoHugging Face18nc1708 /sn38-set-atextn<1K0 likes1.7k downloads11d agoHugging Face19nc1708 /sn38-set-btextn<1K0 likes1.7k downloads11d agoHugging Face20nc1708 /sn38-set-etextn<1K0 likes1.7k downloads11d agoHugging Face21nc1708 /sn38-set-dtextn<1K0 likes1.7k downloads11d agoHugging Face22SetFit /amazon_massive_intent_en-UStext10K<n<100K10 likes1.7k downloads4y agoHugging Face23SetFit /qqp Glue QQP This dataset is a port of the official qqp dataset on the Hub. Note that the question1 and question2 columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M6 likes1.6k downloads5y agoHugging Face24SetFit /mnli Glue MNLI This dataset is a port of the official mnli dataset on the Hub. It contains the matched version. Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively. Also, the test split is not labeled; the label column values are always -1. tabular100K<n<1M8 likes1.5k downloads5y agoHugging Face25marin-dna /genomes-v4-genome_set-animals-intervals-v5_256_128text100M<n<1B0 likes1.5k downloads8mo agoHugging Face26marin-dna /genomes-v5-genome_set-animals-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128 Animals promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 68,286,166 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.text10M<n<100M0 likes1.4k downloads28d agoHugging Face27marin-dna /genomes-v4-genome_set-animals-intervals-v11_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face28marin-dna /genomes-v5-genome_set-animals_order204-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128 204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit main). Each repo in the family is one (genome_set, region-recipe) combination. Size 101,114,252 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.text100M<n<1B0 likes1.3k downloads3mo agoHugging Face29marin-dna /genomes-v4-genome_set-animals-intervals-v10_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face30marin-dna /genomes-v4-genome_set-animals-intervals-v12_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.