datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yahoo_answers_topics
Dataset Card for "Yahoo Answers Topics"
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.Topic-Statementslaws-topics
[!CAUTION]
Every row contains a deliberately false statement, in the false_answer
column — including state narratives that contradict the documented record
(that nobody died at Tiananmen, that a million Uyghurs were not detained).
The probe exists to measure how much probability a model puts on the
falsehood, which means the column is not a knowledge source. This is a
measuring instrument, not training data. Do not fine-tune on it, and if
you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.arXiv_Topics
arXiv Topics Dataset
Dataset Summary
The arXiv Topics Dataset provides a structured mapping of arXiv papers to topic categories at three different levels of abstraction. These topic classifications were generated by prompting GPT-4o, ensuring a hierarchical categorization from broad fields to highly specific research areas.
The dataset consists of 2,422,486 paper IDs, each assigned topics across:
Level 1 (Broad Domains): High-level fields such as Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv_Topics.arXiv-Topics-Embeddings
arXiv Topics Embeddings Dataset
Dataset Summary
The arXiv Topics Embeddings Dataset provides embedding representations for topics associated with arXiv papers. Specifically, the dataset contains embeddings of the arXiv Topics Dataset repository and is used at the retriever module of LitBench to identify relevant papers based on user queries by calculating the similarity between these paper embeddings and the embedding representation of the user query. These… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv-Topics-Embeddings.wikipedia-topicsCreates a pages dataset using Wikipedia.
Explores the 40 root categories and their sub-categories to collect pages. The produced dataset provides up to 2000 pages per category.
See https://github.com/tarekziade/mwcat
yahoo_answers_topicsbeecrowd-beginner-labeled-topics
Beecrowd Beginner Labeled Topics
Dataset Summary
This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.gmat_topics_datasetviet_news_all_topics_1
Dataset Card for "viet_news_all_topics_1"
More Information needed
hausa_voa_topics
Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics)
Dataset Summary
A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Hausa (ISO 639-1: ha)
Dataset Structure
Data Instances
An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.task1592_yahoo_answers_topics_classfication
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1592_yahoo_answers_topics_classfication
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1592_yahoo_answers_topics_classfication.yahoo_answers_topics
Dataset Card for "yahooanswerstopics"
More Information needed
task1594_yahoo_answers_topics_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1594_yahoo_answers_topics_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1594_yahoo_answers_topics_question_generation.fineweb-edu-topics
FineWeb-Edu topic similarities
Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against
the repository's biopsychology, immunopharmacology, and USMLE topic inventories.
Scores are the mean of the top 5 cosine similarities produced by
Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities.
topicsum
Dataset Card for TopicSum Corpus [Single Dataset Comprising of XSUM & DialogSUM for One Liner Summarization/ Topic Generation of Text]
Dataset Description
Links
DialogSUM: https://github.com/cylnlp/dialogsum
XSUM: https://huggingface.co/datasets/knkarthick/xsum
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
TopicSUM is collection of large-scale dialogue summarization dataset from XSUM & DialogSUM, consisting of 241,171… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/topicsum.clinical-trials-trec-topicssensitive-topics-classificationyahoo_answers_topics_sampleThis is a sample from the yahoo_answers_topics dataset. This dataset contains 10% of the original dataset, randomly sampled by class.
Labels follow the following map:
id
label
0
Society & Culture
1
Science & Mathematics
2
Health
3
Education & Reference
4
Computers & Internet
5
Sports
6
Business & Finance
7
Entertainment & Music
8
Family & Relationships
9
Politics & Government
Adaption-video-qa-diverse-topics
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
video_qa_diverse_topics
This dataset contains question-answer pairs derived from a diverse collection of video clips covering topics such as biology, history, sports, and astronomy. Each entry includes a specific question about the video content and a corresponding factual answer, alongside metadata like duration, tags, and object lists. The samples demonstrate a focus on… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-video-qa-diverse-topics.bad-topicsBad Topics is Thai Dataset from Topic Modeling with 2 Bad Topic Website(Bet=16025 , Porn=39237).
task1593_yahoo_answers_topics_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1593_yahoo_answers_topics_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1593_yahoo_answers_topics_classification.corpuslib-topics
CORPUSLIB Topics Dataset
CORPUSLIB — Agentic Corpus Library for Indirect Learning
This dataset contains the topic catalog for DeckerGUI's CORPUSLIB system. CORPUSLIB is a link-gated knowledge library focused on indirect learning as the AI/agentic technology space evolves.
Purpose
Fallback system: When main learning sources are unavailable or undergoing maintenance, CORPUSLIB provides backup topic links
Agent training: Structured topic data for training agentic… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-topics.tech-keywords-topics-summarytext-classification-news-topics
Dataset Card for test
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/test/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/text-classification-news-topics.MSD_manual_topics_user_base
MSD_manual_topics_user_base
This dataset has been built with the website https://www.msdmanuals.com/ provided by Merck & Co for the greater audience.
The MSD manual is an essential source of knowledge for many topics related to symptoms, diseases, health and other related topics. The manual makes an extra effort to make it available both for professionals and patients by having two distinct version.
The content, while being labelled the same, differs by the type of user in order to… See the full description on the dataset page: https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base.turkish_llm_finetune_dataset_4_topics
Turkish LLM Finetune Dataset - 4 Topics
This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System.
Contributors
Barathan Aslan (https://huggingface.co/barathanasln)
Batuhan Kalem(https://huggingface.co/Pancarsuyu)
Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.synthetic-persian-chatbot-topics-retrieval
Dataset Summary
Synthetic Persian Chatbot Topics Retrieval (SynPerChatbotTopicsRetrieval) is a Persian (Farsi) dataset for the Retrieval task, focused on chatbot topic identification. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically created using the GPT-4o-mini language model to simulate user queries and retrieve topic-relevant chatbot responses.
Language(s): Persian (Farsi)
Task(s): Retrieval (Chatbot Topic Retrieval)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-topics-retrieval.chinese-sensitive-topics-qa
Chinese Sensitive Topics QA Dataset
Dataset Summary
This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.topics_labelled
