CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes816 downloads5mo agoHugging Face02owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K27 likes210 downloads4y agoHugging Face03israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes112 downloads4y agoHugging Face04jakeazcona /short-text-multi-labeled-emotion-classificationtabular10K<n<100K2 likes61 downloads5y agoHugging Face05sobamchan /ja-toxic-text-classification-open2ch Open 2ch-based toxic classification dataset Based on p1atdev/open2ch We apply keyword-based filtering to collect toxic texts We use Perspective API to filter non-toxic texts from the original corpus 3k texts for each class, toxic (label=1) and non-toxic (label=0) texts perspective_api_score is a prediction of toxicity score by the Perspective API tabular1K<n<10K1 likes30 downloads2y agoHugging Face06mishrasaurabh847 /covid-tweet-text-classificationtabular10K<n<100K0 likes16 downloads3y agoHugging Face07Nerdy37 /ai-human-text-classification AI vs Human Sentence Classification Dataset Dataset Summary sentence_dataset is a sentence-level binary classification dataset containing approximately 9.84 million sentences labelled as either AI-generated (1) or human-written (0). It was constructed by extracting individual sentences from two source datasets and merging them: Dataset 1 — ai_vs_human_content_v2_20000.csv: 20,000 rows of short text and code snippets with rich metadata (prompt, topic, source… See the full description on the dataset page: https://huggingface.co/datasets/Nerdy37/ai-human-text-classification.tabulartext-classification1M<n<10M0 likes16 downloads3mo agoHugging Face08tommybrenson /text-classificationtabular1K<n<10K0 likes14 downloads2y agoHugging Face09krushilpatel /covid-tweet-text-classificationtabular1K<n<10K0 likes13 downloads3y agoHugging Face10SamagraDataGov /text_classification_testtabularn<1K0 likes6 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.