datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yahoo_answers_topics
Dataset Card for "Yahoo Answers Topics"
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.fineweb-edu-format-topic
FineWeb-Edu w/ Topic and Format Annotations
FineWeb-Edu dataset consists of 1.3T tokens annotated for Topic and Format using wissamantoun/WebOrganizer-TopicClassifier-ModernBERT and wissamantoun/WebOrganizer-FormatClassifier-ModernBERT classifiers.
Similar to WebOrganizer/Corpus-200B but using FineEdu instead of DCLM.
Topic Labels:
Adult
Art & Design
Software Dev.
Crime & Law
Education & Jobs
Hardware
Entertainment
Social Life
Fashion & Beauty
Finance & Business
Food & Dining… See the full description on the dataset page: https://huggingface.co/datasets/wissamantoun/fineweb-edu-format-topic.tweet_topic_multilingual
Dataset Card for "cardiffnlp/tweet_topic_multilingual"
Dataset Summary
This is the official repository of X-Topic (Multilingual Topic Classification in X: Dataset and Analysis, EMNLP 2024), a topic classification dataset based on X (formerly Twitter), featuring 19 topic labels.
The classification task is multi-label, with tweets available in four languages: English, Japanese, Spanish, and Greek.
The dataset comprises 4,000 tweets (1,000 per language), collected between… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multilingual.tweet_topic_multi
Dataset Card for "cardiffnlp/tweet_topic_multi"
Dataset Summary
This is the official repository of TweetTopic ("Twitter Topic Classification
, COLING main conference 2022"), a topic classification dataset on Twitter with 19 labels.
Each instance of TweetTopic comes with a timestamp which distributes from September 2019 to August 2021.
See cardiffnlp/tweet_topic_single for single label version of TweetTopic.
The tweet collection used in TweetTopic is same as what used in… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multi.openalex-topic-title-abstractTopic-Statementsturkish-topic-classification-1.5m
Turkish Topic Classification 1.5M v2
Yirmi konu için anahtar sözcük çeşitlendirmeli Türkçe belge sınıflandırma verisi.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, text, label
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-topic-classification-1.5m.arXiv-Topics-Embeddings
arXiv Topics Embeddings Dataset
Dataset Summary
The arXiv Topics Embeddings Dataset provides embedding representations for topics associated with arXiv papers. Specifically, the dataset contains embeddings of the arXiv Topics Dataset repository and is used at the retriever module of LitBench to identify relevant papers based on user queries by calculating the similarity between these paper embeddings and the embedding representation of the user query. These… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv-Topics-Embeddings.finepdf_fi_edu_score_topic_classifiedopenalex-topic-title-abstractwikipedia-topicsCreates a pages dataset using Wikipedia.
Explores the 40 root categories and their sub-categories to collect pages. The produced dataset provides up to 2000 pages per category.
See https://github.com/tarekziade/mwcat
topic-overwrite
Dataset Card for Topic-Overwrite-Dataset
GitHub | Paper
Summary
This dataset, generated by llava-1.5-7b and labeled by llava-1.6-34b, contains 21k pairs of chosen and rejected answers.
It is used for DPO training in RLHF/RLAIF.
The dataset was created using the processes outlined in the TPO paper, adhering to the Topic-level Preference Overwriting methodology.
It aims to enhance the trustworthiness of MLLM/LVLM and reduce hallucinations.
Usage
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/helehan/topic-overwrite.all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlesyahoo_answers_topicstweet_topic_Llama-3.1-8B-Instruct_vocab_2000_lastFineweb2_fi_edu_score_topic_classifiedall-the-news-2-pythia-tfidf-topic-stratified-v1-articlestask722_mmmlu_answer_generation_random_topic
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task722_mmmlu_answer_generation_random_topic
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task722_mmmlu_answer_generation_random_topic.topic-classification-dataset-realviet_news_all_topics_1
Dataset Card for "viet_news_all_topics_1"
More Information needed
TopicAnnotations-Llama-3.1-8B
WebOrganizer/TopicAnnotations-Llama-3.1-8B
[Paper] [Website] [GitHub]
This dataset contains 1M web pages annotated with topic labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/TopicClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index of the most… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-8B.topic_classification_dataset_genhausa_voa_topics
Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics)
Dataset Summary
A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Hausa (ISO 639-1: ha)
Dataset Structure
Data Instances
An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.task1592_yahoo_answers_topics_classfication
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1592_yahoo_answers_topics_classfication
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1592_yahoo_answers_topics_classfication.yahoo_answers_topics
Dataset Card for "yahooanswerstopics"
More Information needed
hs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.task1594_yahoo_answers_topics_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1594_yahoo_answers_topics_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1594_yahoo_answers_topics_question_generation.fineweb-edu-topics
FineWeb-Edu topic similarities
Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against
the repository's biopsychology, immunopharmacology, and USMLE topic inventories.
Scores are the mean of the top 5 cosine similarities produced by
Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities.
cybersec-topic-classification-dataset
Cybersecurity Topic Classification (CTC) Dataset
Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias.
This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.off-topic
Off-Topic Guardrails Dataset
Overview
This dataset consists of synthetic LLM system prompts paired with user prompts, classified as either off-topic or on-topic. The aim is to provide realistic, real-world-inspired examples reflecting how large language models (LLMs) are used today for both open-ended and closed-ended tasks, such as text generation and classification. This dataset can be used for training and benchmarking off-topic guardrails.
Synthetic Data… See the full description on the dataset page: https://huggingface.co/datasets/gabrielchua/off-topic.
