datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
topic_classificationMultilingual_Topic-Specific_Article-Extraction_and_Classification
Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset
This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text.
Cite the Dataset
Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.Topic_Classification
Dataset Card for News_Topic_Classification
Dataset Description
22462 News Articles classified into 120 different topics
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely article_text and topic.
The article_text column consists of the news article and the topic column consists of the topic each article belongs to
Source Data
The dataset is scrapped from Otherweb database, some news… See the full description on the dataset page: https://huggingface.co/datasets/valurank/Topic_Classification.synthetic-topic-classification-dataset-v1
Tanaos Topic Classification Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.
Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.
Dataset Summary
The dataset contains text samples labeled with their corresponding topics.… See the full description on the dataset page: https://huggingface.co/datasets/donajui/synthetic-topic-classification-dataset-v1.synthetic-topic-classification-dataset-v1
Tanaos Topic Classification Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.
Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.
Dataset Summary
The dataset contains text samples labeled with their corresponding topics.… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-topic-classification-dataset-v1.TOPIC_CLASSIFICATIONTopic-specific-genre-classification_german_historical-newspapers
Dataset Card for Topic-specific Genre Classification of German Historical Newspapers
This dataset was developed to train and evaluate topic-specific genre classification of German-language historical newspaper clippings.
Curated by: [Sarah Oberbichler]
Language(s) (NLP): [German]
License: [afl-3.0]
Uses
Evaluation of machine learning models for topic-specific classification of ocr-processed historical texts with varying quality levels.
Fine-tuning models on… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Topic-specific-genre-classification_german_historical-newspapers.ag-news-topic-classification-processed
AG News Topic Classification Processed Dataset
This dataset is a processed subset of the AG News Classification Dataset from Kaggle.
Task
The task is news topic classification. Each example contains a news title and description combined into a single text field. The goal is to classify each news text into one of four categories:
World
Sports
Business
Sci/Tech
Data Source
The original data comes from the AG News Classification Dataset on Kaggle. For this… See the full description on the dataset page: https://huggingface.co/datasets/Thisisruiii/ag-news-topic-classification-processed.
