datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jigsaw-toxic-comment-classification-challenge
Dataset Description
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
You must create a model which predicts a probability of each type of toxicity for each comment.
File descriptions
train.csv - the training set, contains comments with their binary labels
test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular-benchmark-797-classificationcis5300-text-classification
Complex Word Identification (CIS 5300)
Dataset Description
This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple.
CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.toxicity_classification_jigsaw
Dataset info
Training Dataset:
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
The original dataset can be found here: jigsaw_toxic_classification
Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes.
Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.topic_classificationPubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.ahsan81_hotel-reservations-classification-dataset
Hotel Reservations Dataset
Can you predict if customer is going to cancel the reservation ?
Dataset Info
Source: Kaggle
Original Size: 0.47 MB
Kaggle Downloads: 57,080
Files: 1
Files
Hotel Reservations.csv
Mirrored from Kaggle
llm-classification-distilled-v2-sharded
LLM Classification Distilled v2 Sharded
Overview
This repository stores shard CSV files produced by the teacher-judge distillation pipeline.
How to Use
Run the distillation notebook once per shard:
NUM_SHARDS = 4
SHARD_INDEX = 0 .. 3
After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos.
Final Repositories
Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.Amharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.forecastability_classificationThis dataset is composed of Claude-labelled fineweb documents.
For each document, Claude is asked if it is 'forecastable' (i.e. would be a reasonable seed for a pastcasting question) and to estimate the date the document was published.
V1 splits were generated by having Claude label ~50K random fineweb documents and v2 splits were augmented with labels on ~30K additional documents that a DebertaV3 classifier finetuned on ratio10_v1 thought were forecastable (Claude thought ~1/3 of these… See the full description on the dataset page: https://huggingface.co/datasets/noanabeshima/forecastability_classification.url-classifications
Model Card: URL Classifications Dataset
Dataset Summary
The URL Classifications Dataset is a collection of URL classifications for PDF documents, primarily derived from the SafeDocs corpus. It contains multiple CSV files with different subsets of classifications, including both raw and processed data.
Supported Tasks
This dataset supports the following tasks:
Text Classification
URL-based Document Classification
PDF Content Inference
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/snats/url-classifications.long-covid-classification-data
Data Description
Long-COVID related articles have been manually collected by information specialists.Please find further information here.
Size
Training
Development
Test
Total
Positive Examples
215
76
70
345
Negative Examples
199
62
68
345
Total
414
238
138
690
Citation
@article{10.1093/database/baac048,author = {Langnickel, Lisa and Darms, Johannes and Heldt, Katharina and Ducks, Denise and Fluck, Juliane},title = "{Continuous development… See the full description on the dataset page: https://huggingface.co/datasets/llangnickel/long-covid-classification-data.azerbaijani_review_sentiment_classificationAzerbaijani Sentiment Classification Dataset with ~160K reviews.
Dataset contains 3 columns: Content, Score, Upvotes
short-text-multi-labeled-emotion-classificationNear-Earth-Comets-Classification-dataset
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me
Near-Earth Comets (NECs) Classification Dataset
This repository provides an open-source datasetof Near‑Earth Comets
(NECs) and their classification as Potentially Hazardous… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/Near-Earth-Comets-Classification-dataset.International_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024trustpilot_review_classification
Github Repo: https://github.com/pattplatt/trustpilot-review-classification
Project report: https://homepageblob.blob.core.windows.net/content/Using NLP Techniques to Infer Trustpilot Ratings from User Reviews.pdf
crypto-tax-classification-2026h1
Crypto tax transaction classification, 2026 H1 labelled benchmark
388 public blockchain transactions on Ethereum, Polygon, and Arbitrum, each
labelled with a transaction category and the default US tax treatment that
category maps to under an open, versioned classification schema. Published under
CC BY 4.0 together with the schema itself.
Canonical landing page: https://cryptotaxedge.com/research/benchmarks/2026-h1/
Schema: https://cryptotaxedge.com/standard/
DOI:… See the full description on the dataset page: https://huggingface.co/datasets/CryptoTaxEdge/crypto-tax-classification-2026h1.Sinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK,
Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.traceix-synthetic-classification-data
Traceix Synthetic Generated Data
Traceix Synthetic Generated Data is a synthetic tabular dataset for cybersecurity machine learning research, focused on static Windows PE-file metadata and binary classification workflows.
The dataset contains synthetically generated feature rows designed to resemble metadata patterns commonly extracted from Windows executable files during static analysis. It is intended for experimentation with malware/safe classification, anomaly detection, model… See the full description on the dataset page: https://huggingface.co/datasets/PerkinsFund/traceix-synthetic-classification-data.waste-classificationFake-News-ClassificationDevelop a machine learning program to identify when an article might be fake news. Run by the UTK Machine Learning Club.
This is the Dataset to the Fake-News-Classifier competition in Kaggle. There is a Test csv to check for predictions.
Citation
William Lifferth. (2018). Fake News. Kaggle. https://kaggle.com/competitions/fake-news
ja-toxic-text-classification-open2ch
Open 2ch-based toxic classification dataset
Based on p1atdev/open2ch
We apply keyword-based filtering to collect toxic texts
We use Perspective API to filter non-toxic texts from the original corpus
3k texts for each class, toxic (label=1) and non-toxic (label=0) texts
perspective_api_score is a prediction of toxicity score by the Perspective API
chronos-historical-dataset-sdt-phase-classificationThe schema o fthe dataset is the following:
polityid: a Polity ID formatted with a standard method: 2 letters to indicate the area of origin of the culture, 3 letters to indicate the name of the polity, 1 letter to indicate the type of society (c=culture/community; n=nomads; e=empire; k=kingdom; r=republic) and 1 letter to indicate the periodization (t=terminal; l=late; m=middle; e=early; f=formative; i=initial; =any). For example “EsSpael” is the late Spanish Empire, “ItRomre” is the early… See the full description on the dataset page: https://huggingface.co/datasets/facells/chronos-historical-dataset-sdt-phase-classification.waste-classification-v2
Dataset Card for Dataset Name
Dataset Summary
Dataset used to train a language model to do classification on 50 different waste classes.
Languages
English
Dataset Structure
Data Instances
Phrase
Class
Index
"I have this apple phone charger to throw, where should I put it ?"
PHONE CHARGER
26
"Should I recycle a disposable cup ?"
Plastic Cup
32
"I have a milk brick"
Tetrapack
45
Data Fields
Phrase
Class… See the full description on the dataset page: https://huggingface.co/datasets/thomasavare/waste-classification-v2.emergency_classification
Emergency Messages Classification Dataset
Mushroom_Attribute_Classificationidiom_classificationforecastability_classification_oldtraffy-fondue-organization-classification
