datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cis5300-text-classification
Complex Word Identification (CIS 5300)
Dataset Description
This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple.
CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.stackoverflow-unified-text-open-status-classification
Dataset Card for "stackoverflow-unified-text-open-status-classification"
More Information needed
Amharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.short-text-multi-labeled-emotion-classificationMental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/ihsansaad24/Mental-Health_Text-Classification_Dataset.ai-generated-text-classification
Dataset Card for "ai-generated-text-classification"
More Information needed
stackoverflow-unified-text-open-status-classification-sample
Dataset Card for "stackoverflow-open-status-classification"
More Information needed
Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/Sam20032212/Mental-Health_Text-Classification_Dataset.ja-toxic-text-classification-open2ch
Open 2ch-based toxic classification dataset
Based on p1atdev/open2ch
We apply keyword-based filtering to collect toxic texts
We use Perspective API to filter non-toxic texts from the original corpus
3k texts for each class, toxic (label=1) and non-toxic (label=0) texts
perspective_api_score is a prediction of toxicity score by the Perspective API
hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512
Dataset Card for "hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512"
More Information needed
text-classification-comparison
Text Classification Comparison: Supervised vs Unsupervised on stanfordnlp/imdb
Dataset
stanfordnlp/imdb (Maas et al., 2011)
50,000 IMDB movie reviews: 25,000 train / 25,000 test
Binary sentiment: 0 (negative) / 1 (positive), perfectly balanced in both splits
Preprocessing: TF-IDF (15,000 features, bigrams, sublinear TF, min_df=3)
Models Chosen
Model
Type
Key Reference
Why
Logistic Regression
Linear supervised
McFadden (1974); Ng &… See the full description on the dataset page: https://huggingface.co/datasets/IntimateUser6969/text-classification-comparison.Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/Thamizhmani07/Mental-Health_Text-Classification_Dataset.Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/nayma11/Mental-Health_Text-Classification_Dataset.covid-tweet-text-classificationai-human-text-classification
AI vs Human Sentence Classification Dataset
Dataset Summary
sentence_dataset is a sentence-level binary classification dataset containing approximately 9.84 million sentences labelled as either AI-generated (1) or human-written (0).
It was constructed by extracting individual sentences from two source datasets and merging them:
Dataset 1 — ai_vs_human_content_v2_20000.csv: 20,000 rows of short text and code snippets with rich metadata (prompt, topic, source… See the full description on the dataset page: https://huggingface.co/datasets/Nerdy37/ai-human-text-classification.text-classificationMental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/theaadityapaul/Mental-Health_Text-Classification_Dataset.covid-tweet-text-classificationartificial_text_classification
Artificial Text Classification Dataset
Dataset Summary
The Artificial Text Classification dataset is designed to distinguish between human-generated and machine-generated text. This dataset provides labeled examples of text, enabling researchers and developers to train and evaluate machine learning models for text classification tasks.
Key features:
Text samples: Includes both human-written and machine-generated text.
Labels: Binary target variable where:
1 =… See the full description on the dataset page: https://huggingface.co/datasets/ds-claudia/artificial_text_classification.text-classification-checkpoint-downloadsMental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/balgabekj/Mental-Health_Text-Classification_Dataset.text_classification_testoutput_of_text_classification
