CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zeroshot /twitter-financial-news-topic Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic. The dataset holds 21,107 documents annotated with 20 labels: topics = { "LABEL_0": "Analyst Update", "LABEL_1": "Fed | Central Banks", "LABEL_2": "Company | Product News", "LABEL_3": "Treasuries | Corporate Debt", "LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.texttext-classification10K<n<100K43 likes1.8k downloads3y agoHugging Face02jonaskoenig /topic_classificationtabular10M<n<100M1 likes244 downloads4y agoHugging Face03gvic-unb /beecrowd-beginner-labeled-topics Beecrowd Beginner Labeled Topics Dataset Summary This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.tabulartext-classificationn<1K0 likes208 downloads2mo agoHugging Face04TopicNet /20-Newsgroups 20 Newsgroups Train Some measurable characteristics of the dataset: D — number of documents W — modality dictionary size (number of unique tokens) len D — average document length in modality tokens (number of tokens) len D uniq — average document length in unique modality tokens (number of unique tokens) D @lemmatized W @lemmatized len D @lemmatized len D uniq @bigram W @bigram len D @bigram len D uniq value 11301 1.0614e+06 93.9204 60.5687 213701 18.9099… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/20-Newsgroups.tabulartext-classification10K<n<100K1 likes148 downloads2y agoHugging Face05knkarthick /topicsum Dataset Card for TopicSum Corpus [Single Dataset Comprising of XSUM & DialogSUM for One Liner Summarization/ Topic Generation of Text] Dataset Description Links DialogSUM: https://github.com/cylnlp/dialogsum XSUM: https://huggingface.co/datasets/knkarthick/xsum Point of Contact: https://huggingface.co/knkarthick Dataset Summary TopicSUM is collection of large-scale dialogue summarization dataset from XSUM & DialogSUM, consisting of 241,171… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/topicsum.textsummarization100K<n<1M8 likes146 downloads4y agoHugging Face06Javtor /biomedical-topic-categorizationtext1M<n<10M0 likes110 downloads4y agoHugging Face07TopicNet /RTL-Wiki RTL-Wiki Some measurable characteristics of the dataset: D — number of documents W — modality dictionary size (number of unique tokens) len D — average document length in modality tokens (number of tokens) len D uniq — average document length in unique modality tokens (number of unique tokens) D @lemmatized W @lemmatized len D @lemmatized len D uniq @bigram W @bigram len D @bigram len D uniq value 7838 1.28065e+07 1633.9 691.157 503619 64.2535 30.8372… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/RTL-Wiki.texttext-classification1K<n<10K0 likes65 downloads3y agoHugging Face08barathanasln /turkish_llm_finetune_dataset_4_topics Turkish LLM Finetune Dataset - 4 Topics This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System. Contributors Barathan Aslan (https://huggingface.co/barathanasln) Batuhan Kalem(https://huggingface.co/Pancarsuyu) Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.texttable-question-answering10K<n<100K11 likes61 downloads2y agoHugging Face09jennylester /extreme-topic-fb435f extreme-topic-fb435f Synthetic sensors test data: 50 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/jennylester/extreme-topic-fb435f.tabularn<1K0 likes61 downloads11d agoHugging Face10oberbics /Multilingual_Topic-Specific_Article-Extraction_and_Classification Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text. Cite the Dataset Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.texttext-classificationn<1K1 likes49 downloads1y agoHugging Face11mahfuzh74 /climate-topictext1K<n<10K0 likes48 downloads2y agoHugging Face12TopicNet /PostNauka PostNauka Some measurable characteristics of the dataset: D — number of documents W — modality dictionary size (number of unique tokens) len D — average document length in modality tokens (number of tokens) len D uniq — average document length in unique modality tokens (number of unique tokens) D @title W @title len D @title len D uniq @2gramm W @2gramm len D @2gramm len D uniq @3gramm W @3gramm len D @3gramm len D uniq @snippet W @snippet len D @snippet len D uniq @word… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/PostNauka.texttext-classification1K<n<10K1 likes46 downloads2y agoHugging Face13CharlesBon /chinese-sensitive-topics-qa Chinese Sensitive Topics QA Dataset Dataset Summary This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.textquestion-answeringn<1K2 likes43 downloads9mo agoHugging Face14valurank /Topic_Classification Dataset Card for News_Topic_Classification Dataset Description 22462 News Articles classified into 120 different topics Languages The text in the dataset is in English Dataset Structure The dataset consists of two columns namely article_text and topic. The article_text column consists of the news article and the topic column consists of the topic each article belongs to Source Data The dataset is scrapped from Otherweb database, some news… See the full description on the dataset page: https://huggingface.co/datasets/valurank/Topic_Classification.texttext-classification10K<n<100K4 likes38 downloads3y agoHugging Face15dstefa /New_York_Times_Topicstext100K<n<1M3 likes36 downloads3y agoHugging Face16mtek2000 /hausa_newsclass_topictext1K<n<10K0 likes36 downloads3y agoHugging Face17hajili /azsci_topics Dataset Information This dataset contains titles, topics and subtopics of dissertations written at Azerbaijani universities and institutes. texttext-classification1K<n<10K1 likes36 downloads3y agoHugging Face18Tadesse /news-topictext10K<n<100K0 likes35 downloads3y agoHugging Face19TopicNet /Brown Brown The Brown Corpus was the first million-word electronic corpus of English, created in 1961 at Brown University. This corpus contains text from 500 sources, and the sources have been categorized by genre, such as news, editorial, and so on. 1.1 gives an example of each genre (for a complete list, see http://icame.uib.no/brown/bcm-los.html). Language: English Number of topics: 15 Number of articles: 500 Year: 1961 References NLTK… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/Brown.texttext-classification1K<n<10K0 likes35 downloads2y agoHugging Face20taichisato /hot-topic-be618a hot-topic-be618a Synthetic products test data: 55 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/taichisato/hot-topic-be618a.tabularn<1K0 likes34 downloads11d agoHugging Face21NiGuLa /Russian_Sensitive_Topics General concept of the model Sensitive topics are such topics that have a high chance of initiating a toxic conversation: homophobia, politics, racism, etc. This dataset uses 18 topics. More details can be found in this article presented at the workshop for Balto-Slavic NLP at the EACL-2021 conference. This paper presents the first version of this dataset. Here you can see the last version of the dataset which is significantly larger and also properly filtered.… See the full description on the dataset page: https://huggingface.co/datasets/NiGuLa/Russian_Sensitive_Topics.tabulartext-classification10K<n<100K14 likes33 downloads2y agoHugging Face22Gopher-Lab /X_Twitter_Trending_Topics_August2025 🐦 X-Twitter Scraper: Real-Time Search and Data Extraction Tool Search and scrape X-Twitter (formerly Twitter) for posts by keyword, account, or trending topics.This no-code tool makes it easy to generate real-time, LLM-ready datasets for any AI or content use case. Get started with real-time scraping and instantly structure tweet data into clean JSON. Start Scraping 🚀 Key Features ⚡ Real-Time Fetch – Stream the latest tweets the moment they’re posted 🎯 Flexible… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/X_Twitter_Trending_Topics_August2025.textfeature-extraction10K<n<100K1 likes32 downloads1y agoHugging Face23TopicNet /WikiRef-220 WikiRef220 References Gialampoukidis, I., Vrochidis, S., & Kompatsiaris, I. (2016). A Hybrid Framework for News Clustering Based on the DBSCAN-Martingale and LDA. In Machine Learning and Data Mining in Pattern Recognition (pp. 170-184). Springer International Publishing. texttext-classificationn<1K0 likes31 downloads3y agoHugging Face24TopicNet /ICD-10 ICD-10 (МКБ-10) Some measurable characteristics of the dataset: D — number of documents W — modality dictionary size (number of unique tokens) len D — average document length in modality tokens (number of tokens) len D uniq — average document length in unique modality tokens (number of unique tokens) D @text W @text len D @text len D uniq @letter W @letter len D @letter len D uniq value 1733 953168 550.01 550.01 1733 1 1 Information about document lengths in… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/ICD-10.texttext-classification1K<n<10K0 likes29 downloads3y agoHugging Face25TopicNet /Reuters Reuters The Reuters Corpus contains 10,788 news documents totaling 1.3 million words. The documents have been classified into 90 topics, and grouped into two sets, called "training" and "test"; thus, the text with fileid 'test/14826' is a document drawn from the test set. This split is for training and testing algorithms that automatically detect the topic of a document, as we will see in chap-data-intensive. Language: English Number of topics: 90 Number of articles:… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/Reuters.texttext-classification10K<n<100K0 likes29 downloads2y agoHugging Face26AgileAndy /Professional_and_Hobby_Topicstext100K<n<1M0 likes29 downloads1y agoHugging Face27jensjorisdecorte /TU-Expert-Collection-Topic-Synonymstext1K<n<10K0 likes28 downloads2y agoHugging Face28AmanPriyanshu /Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens) 📝Check out the Blog Post This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000 samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing. Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.textsummarization100K<n<1M8 likes28 downloads2y agoHugging Face29donajui /synthetic-topic-classification-dataset-v1 Tanaos Topic Classification Training Dataset This dataset was created synthetically by Tanaos with the Artifex Python library. The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories. Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset. Dataset Summary The dataset contains text samples labeled with their corresponding topics.… See the full description on the dataset page: https://huggingface.co/datasets/donajui/synthetic-topic-classification-dataset-v1.texttext-classification10K<n<100K0 likes27 downloads7mo agoHugging Face30Javtor /biomedical-topic-categorization-validationtext100K<n<1M1 likes25 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.