datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Topic-Statementstweet_topic_Llama-3.1-8B-Instruct_vocab_2000_lastfinepdf_fi_edu_score_topic_classifiedFineweb2_fi_edu_score_topic_classifiedyahoo_answers_topics
Dataset Card for "yahooanswerstopics"
More Information needed
TopicAnnotations-Llama-3.1-8B
WebOrganizer/TopicAnnotations-Llama-3.1-8B
[Paper] [Website] [GitHub]
This dataset contains 1M web pages annotated with topic labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/TopicClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index of the most… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-8B.all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlesall-the-news-2-pythia-tfidf-topic-stratified-v1-articleshs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.cybersec-topic-classification-dataset
Cybersecurity Topic Classification (CTC) Dataset
Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias.
This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.fineweb-edu-topics
FineWeb-Edu topic similarities
Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against
the repository's biopsychology, immunopharmacology, and USMLE topic inventories.
Scores are the mean of the top 5 cosine similarities produced by
Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities.
off-topic
Off-Topic Guardrails Dataset
Overview
This dataset consists of synthetic LLM system prompts paired with user prompts, classified as either off-topic or on-topic. The aim is to provide realistic, real-world-inspired examples reflecting how large language models (LLMs) are used today for both open-ended and closed-ended tasks, such as text generation and classification. This dataset can be used for training and benchmarking off-topic guardrails.
Synthetic Data… See the full description on the dataset page: https://huggingface.co/datasets/gabrielchua/off-topic.clinical-trials-trec-topicsbiomedical-topic-categorization-cased
Dataset Card for "biomedical-topic-categorization-cased"
More Information needed
topic_based_nli_test
Dataset Card for "test_topicbasednli"
More Information needed
corral-QAs-topic_reports
Corral – QA Topic Reports
Averaged QA results for factual-knowledge and reasoning evaluations across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments.
The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports.all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencesTopicAnnotations-Llama-3.1-405B-FP8
WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8
[Paper] [Website] [GitHub]
This dataset contains 100K web pages annotated with topic labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/TopicClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8.tweet_topic_ERNIE-4.5-0.3B-PT_vocab_2000_lasttech-keywords-topics-summarytweet_topic_ERNIE-4.5-0.3B-PT_vocab_4000_lastwikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-articlestopic-classificationtweet_topic_ERNIE-4.5-0.3B-PT_vocab_1000_lasttopics_labelledtweet_topic_ERNIE-4.5-0.3B-PT_vocab_500_lastall-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencestweet_topic_Llama-3.2-1B-Instruct_vocab_2000_lastcup_it_ds_split_with_lang_with_topic
Dataset Card for "cup_it_ds_split_with_lang_with_topic"
More Information needed
wikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-articles
