datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Topic-Statementsall-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlesall-the-news-2-pythia-tfidf-topic-stratified-v1-articlestopic_classificationfinepdf_fi_edu_score_topic_classifiedtweet_topic_Llama-3.1-8B-Instruct_vocab_2000_lastFineweb2_fi_edu_score_topic_classifiedbeecrowd-beginner-labeled-topics
Beecrowd Beginner Labeled Topics
Dataset Summary
This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencesTopicAnnotations-Llama-3.1-8B
WebOrganizer/TopicAnnotations-Llama-3.1-8B
[Paper] [Website] [GitHub]
This dataset contains 1M web pages annotated with topic labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/TopicClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index of the most… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-8B.yahoo_answers_topics
Dataset Card for "yahooanswerstopics"
More Information needed
hs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.all-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequences20-Newsgroups
20 Newsgroups
Train
Some measurable characteristics of the dataset:
D — number of documents
W — modality dictionary size (number of unique tokens)
len D — average document length in modality tokens (number of tokens)
len D uniq — average document length in unique modality tokens (number of unique tokens)
D
@lemmatized W
@lemmatized len D
@lemmatized len D uniq
@bigram W
@bigram len D
@bigram len D uniq
value
11301
1.0614e+06
93.9204
60.5687
213701
18.9099… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/20-Newsgroups.fineweb-edu-topics
FineWeb-Edu topic similarities
Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against
the repository's biopsychology, immunopharmacology, and USMLE topic inventories.
Scores are the mean of the top 5 cosine similarities produced by
Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities.
cybersec-topic-classification-dataset
Cybersecurity Topic Classification (CTC) Dataset
Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias.
This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.off-topic
Off-Topic Guardrails Dataset
Overview
This dataset consists of synthetic LLM system prompts paired with user prompts, classified as either off-topic or on-topic. The aim is to provide realistic, real-world-inspired examples reflecting how large language models (LLMs) are used today for both open-ended and closed-ended tasks, such as text generation and classification. This dataset can be used for training and benchmarking off-topic guardrails.
Synthetic Data… See the full description on the dataset page: https://huggingface.co/datasets/gabrielchua/off-topic.clinical-trials-trec-topicswikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-articlestopic_based_nli_test
Dataset Card for "test_topicbasednli"
More Information needed
biomedical-topic-categorization-cased
Dataset Card for "biomedical-topic-categorization-cased"
More Information needed
corral-QAs-topic_reports
Corral – QA Topic Reports
Averaged QA results for factual-knowledge and reasoning evaluations across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments.
The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports.tech-keywords-topics-summaryTopicAnnotations-Llama-3.1-405B-FP8
WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8
[Paper] [Website] [GitHub]
This dataset contains 100K web pages annotated with topic labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/TopicClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8.topic-classificationwikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-articlesextreme-topic-fb435f
extreme-topic-fb435f
Synthetic sensors test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/jennylester/extreme-topic-fb435f.tweet_topic_ERNIE-4.5-0.3B-PT_vocab_4000_lastwikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-val-sequencestweet_topic_ERNIE-4.5-0.3B-PT_vocab_2000_last
