datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biomedical-topic-categorizationNews_Articles_Categorization
Dataset Card for News_Articles_Categorization
Dataset Description
3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely Text and Category.
The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.indian-transaction-categorization-synthetic
Synthetic Indian Bank Transaction Narrations
810 synthetic (text, category) pairs mimicking Indian bank/credit-card statement narrations —
built to train the Sumeetgpt/indian-transaction-categorizer
SetFit model.
Why this exists
While building a personal finance app, we searched for a public dataset pairing real Indian
transaction narration formats (UPI, NEFT, IMPS, ACH) with spending-category labels, and found
none: datasets with real-looking Indian narration… See the full description on the dataset page: https://huggingface.co/datasets/Sumeetgpt/indian-transaction-categorization-synthetic.scientific-paragraphs-categorization
A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific
We present a dataset of 833k paragraphs extracted from CC-BY licensed
scientific publications, classified into four categories: acknowledgments, data
mentions, software/code mentions, and clinical trial mentions. The paragraphs
are primarily in English and French, with additional European languages
represented. Each paragraph is annotated with language identification (using
fastText) and scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/dataesr/scientific-paragraphs-categorization.News-Article-Categorization_IAB
Article and Category Dataset
Overview
This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more.
Dataset Information
Number of Samples: 871,909
Number of Categories: 26
Column Information
text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.biomedical-topic-categorization-validationbank-statement-categorizationsupport-message-categorizationarabic_categorization_dataResume_Categorizationtext-simple-categorizationcategorization_data
