CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thesofakillers /jigsaw-toxic-comment-classification-challenge Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate You must create a model which predicts a probability of each type of toxicity for each comment. File descriptions train.csv - the training set, contains comments with their binary labels test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular100K<n<1M13 likes15k downloads2y agoHugging Face02jackhhao /jailbreak-classification Jailbreak Classification Dataset Summary Dataset used to classify prompts as jailbreak vs. benign. Dataset Structure Data Fields prompt: an LLM prompt type: classification label, either jailbreak or benign Dataset Creation Curation Rationale Created to help detect & prevent harmful jailbreak prompts when users interact with LLMs. Source Data Jailbreak prompts sourced from: https://github.com/verazuo/jailbreak_llms… See the full description on the dataset page: https://huggingface.co/datasets/jackhhao/jailbreak-classification.texttext-classification1K<n<10K83 likes3.1k downloads3y agoHugging Face03meriemm6 /commit-classification-dataset Commit Classification Dataset This dataset is designed for multi-label classification of Git commit messages into predefined categories. Dataset Summary This dataset contains: Training data: Commit messages and their corresponding labels for training the model. Validation data: A separate set of messages for tuning and evaluation. Testing data: Unlabeled commit messages for testing the model’s performance. The goal of the dataset is to classify each commit message into… See the full description on the dataset page: https://huggingface.co/datasets/meriemm6/commit-classification-dataset.texttext-classification1K<n<10K0 likes1.5k downloads2y agoHugging Face04rsh-raj /commit-classification-17ktext10K<n<100K0 likes1.1k downloads2y agoHugging Face05imodels /tabular-benchmark-797-classificationtabular1K<n<10K0 likes1k downloads3y agoHugging Face06CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes816 downloads5mo agoHugging Face07Kushal0532 /news-political-bias-classification-datasetDataset actually from kaggle. Couldn't find it here so I uploaded it. text10K<n<100K0 likes755 downloads11mo agoHugging Face08seanswyi /sms-spam-classificationtext1K<n<10K0 likes624 downloads2y agoHugging Face09Arsive /toxicity_classification_jigsaw Dataset info Training Dataset: You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate The original dataset can be found here: jigsaw_toxic_classification Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes. Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.tabulartext-classification100K<n<1M5 likes392 downloads3y agoHugging Face10imanoop7 /phishing_url_classification Phishing URL Classification Dataset This dataset contains URLs labeled as 'Safe' (0) or 'Not Safe' (1) for phishing detection tasks. Dataset Summary This dataset contains URLs labeled for phishing detection tasks. It's designed to help train and evaluate models that can identify potentially malicious URLs. Dataset Creation The dataset was synthetically generated using a custom script that creates both legitimate and potentially phishing URLs. This approach… See the full description on the dataset page: https://huggingface.co/datasets/imanoop7/phishing_url_classification.texttext-classification100K<n<1M5 likes276 downloads2y agoHugging Face11samscript18 /adaption-defi-wallet-risk-classification-v1 This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-defi_wallet_risk_classification This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks based on behavioral features. Each sample provides wallet metrics such as transaction counts, action ratios, and concentration levels, followed by a binary risk label and a concise reasoning statement. The data is designed for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification-v1.text1K<n<10K0 likes270 downloads3mo agoHugging Face12jonaskoenig /topic_classificationtabular10M<n<100M0 likes242 downloads4y agoHugging Face13bhargavi909 /cancer_data_classificationtext1K<n<10K0 likes228 downloads3y agoHugging Face14owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K27 likes210 downloads4y agoHugging Face15jason1966 /ahsan81_hotel-reservations-classification-dataset Hotel Reservations Dataset Can you predict if customer is going to cancel the reservation ? Dataset Info Source: Kaggle Original Size: 0.47 MB Kaggle Downloads: 57,080 Files: 1 Files Hotel Reservations.csv Mirrored from Kaggle tabular10K<n<100K0 likes189 downloads6mo agoHugging Face16star092304 /typhoon-intensity-classification Typhoon - Image Classification Dataset This dataset comes from PTIT AI Challenge and is organized for a multi-class image classification task focusing on tropical cyclone (typhoon) intensity estimation. Dataset Structure The directory structure is organized as follows: train/ ├── images/ │ ├── image1.jpg │ └── ... └── annotations.csv (only present in the train folder) The public_test and private_test sets are used to evaluate and score the… See the full description on the dataset page: https://huggingface.co/datasets/star092304/typhoon-intensity-classification.imageimage-classification1K<n<10K1 likes180 downloads3mo agoHugging Face17knowledgator /Scientific-text-classificationtext10K<n<100K16 likes161 downloads3y agoHugging Face18pgurazada1 /patent_classificationtextn<1K1 likes154 downloads2y agoHugging Face19UniqueData /email-spam-classification Email Spam Classification The dataset consists of a collection of emails categorized into two major classes: spam and not spam. It is designed to facilitate the development and evaluation of spam detection or email filtering systems. The spam emails in the dataset are typically unsolicited and unwanted messages that aim to promote products or services, spread malware, or deceive recipients for various malicious purposes. These emails often contain misleading subject lines… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/email-spam-classification.texttext-classificationn<1K10 likes122 downloads1y agoHugging Face20tussiiiii /llm-classification-distilled-v2-sharded LLM Classification Distilled v2 Sharded Overview This repository stores shard CSV files produced by the teacher-judge distillation pipeline. How to Use Run the distillation notebook once per shard: NUM_SHARDS = 4 SHARD_INDEX = 0 .. 3 After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos. Final Repositories Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.tabulartext-classification100K<n<1M0 likes122 downloads4mo agoHugging Face21bhargavi909 /cancer_classificationtext1K<n<10K0 likes121 downloads3y agoHugging Face22ml4pubmed /pubmed-classification-20k ml4pubmed/pubmed-classification-20k 20k subset of pubmed text classification from course texttext-classification100K<n<1M1 likes118 downloads4y agoHugging Face23israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes112 downloads4y agoHugging Face24zm-hf /dianping-classificationtexttext-classification10K<n<100K1 likes107 downloads2y agoHugging Face25ctoraman /BilCat-news-classificationBilCat: Bilkent Text Classification (News Categorization) Dataset 7540 Turkish news articles (Milliyet and TRT merged) with category labels (Dunya, Ekonomi, Politika, KulturSanat, Saglik, Spor, Turkiye, Yazarlar). Column header is the first line. Other details are at https://github.com/BilkentInformationRetrievalGroup/BilCat/ Citation: C. Toraman, F. Can and S. Koçberber. Developing a text categorization template for Turkish news portals. 2011 International Symposium on Innovations in… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/BilCat-news-classification.texttext-classification1K<n<10K0 likes97 downloads3y agoHugging Face26will4381 /job-posting-classificationOriginal job descriptions were derived from xanderios/linkedin-job-postings, and classification data was created synthetically with GPT-4o-Mini. All values not represented nor found in the job description are marked as null. Note: Some responses are hallucinations, despite maintaining the correct .json format, the content is wrong. All instances of incorrect .json formatting have been removed from the dataset, hallucinated content however still remains. Future: Might consider sourcing a resume… See the full description on the dataset page: https://huggingface.co/datasets/will4381/job-posting-classification.text10K<n<100K7 likes94 downloads2y agoHugging Face27noanabeshima /forecastability_classificationThis dataset is composed of Claude-labelled fineweb documents. For each document, Claude is asked if it is 'forecastable' (i.e. would be a reasonable seed for a pastcasting question) and to estimate the date the document was published. V1 splits were generated by having Claude label ~50K random fineweb documents and v2 splits were augmented with labels on ~30K additional documents that a DebertaV3 classifier finetuned on ratio10_v1 thought were forecastable (Claude thought ~1/3 of these… See the full description on the dataset page: https://huggingface.co/datasets/noanabeshima/forecastability_classification.tabular100K<n<1M0 likes92 downloads1y agoHugging Face28mlexplorer008 /malayalam_news_classificationtext1K<n<10K0 likes87 downloads2y agoHugging Face29marksverdhei /clickbait_title_classificationDataset introduced in Stop Clickbait: Detecting and Preventing Clickbaits in Online News Mediaby Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, Niloy Ganguly Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, and Niloy Ganguly. "Stop Clickbait: Detecting and Preventing Clickbaits in Online News Media”. In Proceedings of the 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), San Fransisco, US, August 2016. Cite:… See the full description on the dataset page: https://huggingface.co/datasets/marksverdhei/clickbait_title_classification.text10K<n<100K6 likes85 downloads4y agoHugging Face30snats /url-classifications Model Card: URL Classifications Dataset Dataset Summary The URL Classifications Dataset is a collection of URL classifications for PDF documents, primarily derived from the SafeDocs corpus. It contains multiple CSV files with different subsets of classifications, including both raw and processed data. Supported Tasks This dataset supports the following tasks: Text Classification URL-based Document Classification PDF Content Inference Languages The… See the full description on the dataset page: https://huggingface.co/datasets/snats/url-classifications.tabular1M<n<10M12 likes80 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.