datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter-financial-news-topic
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic.
The dataset holds 21,107 documents annotated with 20 labels:
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.topic_classificationbeecrowd-beginner-labeled-topics
Beecrowd Beginner Labeled Topics
Dataset Summary
This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.20-Newsgroups
20 Newsgroups
Train
Some measurable characteristics of the dataset:
D — number of documents
W — modality dictionary size (number of unique tokens)
len D — average document length in modality tokens (number of tokens)
len D uniq — average document length in unique modality tokens (number of unique tokens)
D
@lemmatized W
@lemmatized len D
@lemmatized len D uniq
@bigram W
@bigram len D
@bigram len D uniq
value
11301
1.0614e+06
93.9204
60.5687
213701
18.9099… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/20-Newsgroups.topicsum
Dataset Card for TopicSum Corpus [Single Dataset Comprising of XSUM & DialogSUM for One Liner Summarization/ Topic Generation of Text]
Dataset Description
Links
DialogSUM: https://github.com/cylnlp/dialogsum
XSUM: https://huggingface.co/datasets/knkarthick/xsum
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
TopicSUM is collection of large-scale dialogue summarization dataset from XSUM & DialogSUM, consisting of 241,171… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/topicsum.biomedical-topic-categorizationRTL-Wiki
RTL-Wiki
Some measurable characteristics of the dataset:
D — number of documents
W — modality dictionary size (number of unique tokens)
len D — average document length in modality tokens (number of tokens)
len D uniq — average document length in unique modality tokens (number of unique tokens)
D
@lemmatized W
@lemmatized len D
@lemmatized len D uniq
@bigram W
@bigram len D
@bigram len D uniq
value
7838
1.28065e+07
1633.9
691.157
503619
64.2535
30.8372… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/RTL-Wiki.turkish_llm_finetune_dataset_4_topics
Turkish LLM Finetune Dataset - 4 Topics
This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System.
Contributors
Barathan Aslan (https://huggingface.co/barathanasln)
Batuhan Kalem(https://huggingface.co/Pancarsuyu)
Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.extreme-topic-fb435f
extreme-topic-fb435f
Synthetic sensors test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/jennylester/extreme-topic-fb435f.Multilingual_Topic-Specific_Article-Extraction_and_Classification
Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset
This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text.
Cite the Dataset
Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.climate-topicPostNauka
PostNauka
Some measurable characteristics of the dataset:
D — number of documents
W — modality dictionary size (number of unique tokens)
len D — average document length in modality tokens (number of tokens)
len D uniq — average document length in unique modality tokens (number of unique tokens)
D
@title W
@title len D
@title len D uniq
@2gramm W
@2gramm len D
@2gramm len D uniq
@3gramm W
@3gramm len D
@3gramm len D uniq
@snippet W
@snippet len D
@snippet len D uniq
@word… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/PostNauka.chinese-sensitive-topics-qa
Chinese Sensitive Topics QA Dataset
Dataset Summary
This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.Topic_Classification
Dataset Card for News_Topic_Classification
Dataset Description
22462 News Articles classified into 120 different topics
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely article_text and topic.
The article_text column consists of the news article and the topic column consists of the topic each article belongs to
Source Data
The dataset is scrapped from Otherweb database, some news… See the full description on the dataset page: https://huggingface.co/datasets/valurank/Topic_Classification.New_York_Times_Topicshausa_newsclass_topicazsci_topics
Dataset Information
This dataset contains titles, topics and subtopics of dissertations written at Azerbaijani universities and institutes.
news-topicBrown
Brown
The Brown Corpus was the first million-word electronic corpus of English, created in 1961 at Brown University. This corpus contains text from 500 sources, and the sources have been categorized by genre, such as news, editorial, and so on. 1.1 gives an example of each genre (for a complete list, see http://icame.uib.no/brown/bcm-los.html).
Language: English
Number of topics: 15
Number of articles: 500
Year: 1961
References
NLTK… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/Brown.hot-topic-be618a
hot-topic-be618a
Synthetic products test data: 55 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/taichisato/hot-topic-be618a.Russian_Sensitive_Topics
General concept of the model
Sensitive topics are such topics that have a high chance of initiating a toxic conversation: homophobia, politics, racism, etc. This dataset uses 18 topics.
More details can be found in this article presented at the workshop for Balto-Slavic NLP at the EACL-2021 conference.
This paper presents the first version of this dataset. Here you can see the last version of the dataset which is significantly larger and also properly filtered.… See the full description on the dataset page: https://huggingface.co/datasets/NiGuLa/Russian_Sensitive_Topics.X_Twitter_Trending_Topics_August2025
🐦 X-Twitter Scraper: Real-Time Search and Data Extraction Tool
Search and scrape X-Twitter (formerly Twitter) for posts by keyword, account, or trending topics.This no-code tool makes it easy to generate real-time, LLM-ready datasets for any AI or content use case.
Get started with real-time scraping and instantly structure tweet data into clean JSON.
Start Scraping
🚀 Key Features
⚡ Real-Time Fetch – Stream the latest tweets the moment they’re posted
🎯 Flexible… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/X_Twitter_Trending_Topics_August2025.WikiRef-220
WikiRef220
References
Gialampoukidis, I., Vrochidis, S., & Kompatsiaris, I. (2016). A Hybrid Framework for News Clustering Based on the DBSCAN-Martingale and LDA. In Machine Learning and Data Mining in Pattern Recognition (pp. 170-184). Springer International Publishing.
ICD-10
ICD-10 (МКБ-10)
Some measurable characteristics of the dataset:
D — number of documents
W — modality dictionary size (number of unique tokens)
len D — average document length in modality tokens (number of tokens)
len D uniq — average document length in unique modality tokens (number of unique tokens)
D
@text W
@text len D
@text len D uniq
@letter W
@letter len D
@letter len D uniq
value
1733
953168
550.01
550.01
1733
1
1
Information about document lengths in… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/ICD-10.Reuters
Reuters
The Reuters Corpus contains 10,788 news documents totaling 1.3 million words. The documents have been classified into 90 topics, and grouped into two sets, called "training" and "test"; thus, the text with fileid 'test/14826' is a document drawn from the test set. This split is for training and testing algorithms that automatically detect the topic of a document, as we will see in chap-data-intensive.
Language: English
Number of topics: 90
Number of articles:… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/Reuters.Professional_and_Hobby_TopicsTU-Expert-Collection-Topic-SynonymsDynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens
Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens)
📝Check out the Blog Post
This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000
samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing.
Dataset Overview
Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.synthetic-topic-classification-dataset-v1
Tanaos Topic Classification Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.
Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.
Dataset Summary
The dataset contains text samples labeled with their corresponding topics.… See the full description on the dataset page: https://huggingface.co/datasets/donajui/synthetic-topic-classification-dataset-v1.biomedical-topic-categorization-validation
