datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-topic-classification-1.5m
Turkish Topic Classification 1.5M v2
Yirmi konu için anahtar sözcük çeşitlendirmeli Türkçe belge sınıflandırma verisi.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, text, label
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-topic-classification-1.5m.topic_classificationen-document-topic-classification
English Document Topic Classification Dataset
English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics.
Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks.
Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation).
Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.topic_classification_dataset_gentopic-classification-dataset-realcybersec-topic-classification-dataset
Cybersecurity Topic Classification (CTC) Dataset
Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias.
This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.task379_agnews_topic_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task379_agnews_topic_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task379_agnews_topic_classification.topic-classification-dataset-real-labeledtopic-classification-dataset-real-2topic-classificationFinance_sentiment_and_topic_classification_Translation_English_to_Spanish_v1Multilingual_Topic-Specific_Article-Extraction_and_Classification
Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset
This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text.
Cite the Dataset
Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.indonesian-financial-topic-classification-datasetTranslated version of https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend",
"LABEL_5": "Earnings",
"LABEL_6": "Energy | Oil",
"LABEL_7": "Financials",
"LABEL_8": "Currencies",
"LABEL_9": "General News | Opinion",
"LABEL_10": "Gold | Metals | Materials"… See the full description on the dataset page: https://huggingface.co/datasets/intanm/indonesian-financial-topic-classification-dataset.Topic_Classification
Dataset Card for News_Topic_Classification
Dataset Description
22462 News Articles classified into 120 different topics
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely article_text and topic.
The article_text column consists of the news article and the topic column consists of the topic each article belongs to
Source Data
The dataset is scrapped from Otherweb database, some news… See the full description on the dataset page: https://huggingface.co/datasets/valurank/Topic_Classification.Finance_sentiment_and_topic_classification_En
original dataset.
https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k
topic-classification-datasetkorean-news-topic-classification
Korean News Topic Classification Dataset (Synthetic)
한국어 뉴스 토픽 분류를 위한 합성 데이터셋입니다.
Dataset Description
이 데이터셋은 자연어처리 실습을 위해 제작된 교육용 합성 데이터셋입니다.
Dataset Summary
언어: 한국어 (Korean)
도메인: 뉴스 헤드라인 스타일
태스크: 4-class 텍스트 분류 (Topic Classification)
생성 방식: 템플릿 기반 합성 데이터 (Synthetic)
Supported Tasks
Text Classification: 주어진 문장을 4개 카테고리 중 하나로 분류
N-to-1 Task: 입력 시퀀스 -> 단일 클래스 레이블
Languages
한국어 (Korean, ko)
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/hugmanskj/korean-news-topic-classification.cybersec-topic-classification-dataset-filteredsynthetic-topic-classification-dataset-v1
Tanaos Topic Classification Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.
Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.
Dataset Summary
The dataset contains text samples labeled with their corresponding topics.… See the full description on the dataset page: https://huggingface.co/datasets/donajui/synthetic-topic-classification-dataset-v1.synthetic-topic-classification-dataset-v1
Tanaos Topic Classification Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.
Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.
Dataset Summary
The dataset contains text samples labeled with their corresponding topics.… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-topic-classification-dataset-v1.TOPIC_CLASSIFICATIONEmakhuwa-News-Topic-ClassificationBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-News-Topic-Classification.Topic-specific-genre-classification_german_historical-newspapers
Dataset Card for Topic-specific Genre Classification of German Historical Newspapers
This dataset was developed to train and evaluate topic-specific genre classification of German-language historical newspaper clippings.
Curated by: [Sarah Oberbichler]
Language(s) (NLP): [German]
License: [afl-3.0]
Uses
Evaluation of machine learning models for topic-specific classification of ocr-processed historical texts with varying quality levels.
Fine-tuning models on… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Topic-specific-genre-classification_german_historical-newspapers.indonesian-financial-topic-classification-datasetTranslated version of https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend",
"LABEL_5": "Earnings",
"LABEL_6": "Energy | Oil",
"LABEL_7": "Financials",
"LABEL_8": "Currencies",
"LABEL_9": "General News | Opinion",
"LABEL_10": "Gold | Metals | Materials"… See the full description on the dataset page: https://huggingface.co/datasets/centidiary/indonesian-financial-topic-classification-dataset.topic-classifier-news-headlines-classification
Dataset Card for "topic-classifier-news-headlines-classification"
More Information needed
russian-it-topic-classification
Russian-it-topic-classification
Датасет подготовлен для конкурсного трека по тематической классификации русскоязычных IT-текстов на платформе All Cups.
1) Разметка данных
Задача: мультиклассовая классификация по 6 классам:
programming
machine_learning
infrastructure
cybersecurity
data
hardware
Каждая запись содержит:
id — уникальный идентификатор вида vk_it_XXXXXXXXX;
text — текстовый фрагмент;
label — целевая метка;
source_title, source_url — служебные поля… See the full description on the dataset page: https://huggingface.co/datasets/commifeez/russian-it-topic-classification.hausa_topic_classificationtl-mirea-messages-topic-classification
RTU MIREA Telegram Channel Messages Topic Classification
The dataset has been created to fine-tune mDeBERTa-v3-base-mnli-xnli model for text classification tasks on Telegram messages with 17 different topics specific to the target of classification.
Overview
This dataset has been parsed from the official Telegram channel of RTU MIREA, using aiogram, and manually labelled based on the topic of a given message.
Preprocessing
Preprocessing of the data included… See the full description on the dataset page: https://huggingface.co/datasets/complicat9d/tl-mirea-messages-topic-classification.ag-news-topic-classification-processed
AG News Topic Classification Processed Dataset
This dataset is a processed subset of the AG News Classification Dataset from Kaggle.
Task
The task is news topic classification. Each example contains a news title and description combined into a single text field. The goal is to classify each news text into one of four categories:
World
Sports
Business
Sci/Tech
Data Source
The original data comes from the AG News Classification Dataset on Kaggle. For this… See the full description on the dataset page: https://huggingface.co/datasets/Thisisruiii/ag-news-topic-classification-processed.text-topic-classification
