CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Javtor /biomedical-topic-categorizationtext1M<n<10M0 likes95 downloads4y agoHugging Face02valurank /News_Articles_Categorization Dataset Card for News_Articles_Categorization Dataset Description 3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science Languages The text in the dataset is in English Dataset Structure The dataset consists of two columns namely Text and Category. The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.texttext-classification1K<n<10K5 likes89 downloads3y agoHugging Face03Sumeetgpt /indian-transaction-categorization-synthetic Synthetic Indian Bank Transaction Narrations 810 synthetic (text, category) pairs mimicking Indian bank/credit-card statement narrations — built to train the Sumeetgpt/indian-transaction-categorizer SetFit model. Why this exists While building a personal finance app, we searched for a public dataset pairing real Indian transaction narration formats (UPI, NEFT, IMPS, ACH) with spending-category labels, and found none: datasets with real-looking Indian narration… See the full description on the dataset page: https://huggingface.co/datasets/Sumeetgpt/indian-transaction-categorization-synthetic.texttext-classificationn<1K0 likes73 downloads25d agoHugging Face04dataesr /scientific-paragraphs-categorization A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are primarily in English and French, with additional European languages represented. Each paragraph is annotated with language identification (using fastText) and scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/dataesr/scientific-paragraphs-categorization.texttext-classification100K<n<1M2 likes36 downloads1y agoHugging Face05shishir-dwi /News-Article-Categorization_IAB Article and Category Dataset Overview This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more. Dataset Information Number of Samples: 871,909 Number of Categories: 26 Column Information text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.texttext-classification100K<n<1M4 likes35 downloads3y agoHugging Face06Javtor /biomedical-topic-categorization-validationtext100K<n<1M1 likes25 downloads4y agoHugging Face07karthiksagarn /bank-statement-categorizationtexttext-classification1K<n<10K2 likes20 downloads1y agoHugging Face08kaustuvkunal /support-message-categorizationtextn<1K0 likes17 downloads1y agoHugging Face09abdalrahmanshahrour /arabic_categorization_datatext10K<n<100K0 likes16 downloads4y agoHugging Face10swamy1418 /Resume_Categorizationtexttext-classification1K<n<10K0 likes9 downloads3y agoHugging Face11AkilanSelvam /text-simple-categorizationtext10K<n<100K0 likes3 downloads2y agoHugging Face12indrapurnayasa /categorization_datatextn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.