CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mbzuai-ugrip-statement-tuning /Topic-Statementstabular100K<n<1M0 likes352 downloads2y agoHugging Face02raymondzmc /tweet_topic_Llama-3.1-8B-Instruct_vocab_2000_lasttabular10K<n<100K0 likes272 downloads9mo agoHugging Face03Finnish-NLP /finepdf_fi_edu_score_topic_classifiedtabular1M<n<10M0 likes248 downloads1y agoHugging Face04Finnish-NLP /Fineweb2_fi_edu_score_topic_classifiedtabular10M<n<100M0 likes227 downloads10mo agoHugging Face05pietrolesci /yahoo_answers_topics Dataset Card for "yahooanswerstopics" More Information needed tabular1M<n<10M0 likes201 downloads3y agoHugging Face06WebOrganizer /TopicAnnotations-Llama-3.1-8B WebOrganizer/TopicAnnotations-Llama-3.1-8B [Paper] [Website] [GitHub] This dataset contains 1M web pages annotated with topic labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/TopicClassifier. Dataset Structure Each example contains the following fields: text: The text content of the web page url: The URL of the web page top_choice_index: Index of the most… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-8B.tabular1M<n<10M1 likes195 downloads2y agoHugging Face07MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes168 downloads1mo agoHugging Face08MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes159 downloads1mo agoHugging Face09model-organisms-for-real /hs3-prompt-pool-topic-judged hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training Prompts only (no completions). Every user prompt in model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated 35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0) for the high-level topic of both quirk families. Why Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.tabulartext-generation100K<n<1M0 likes152 downloads7d agoHugging Face10naufalso /cybersec-topic-classification-dataset Cybersecurity Topic Classification (CTC) Dataset Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias. This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.tabular10M<n<100M0 likes148 downloads2y agoHugging Face11LeoZotos /fineweb-edu-topics FineWeb-Edu topic similarities Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against the repository's biopsychology, immunopharmacology, and USMLE topic inventories. Scores are the mean of the top 5 cosine similarities produced by Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities. tabularfeature-extraction10M<n<100M0 likes143 downloads7d agoHugging Face12gabrielchua /off-topic Off-Topic Guardrails Dataset Overview This dataset consists of synthetic LLM system prompts paired with user prompts, classified as either off-topic or on-topic. The aim is to provide realistic, real-world-inspired examples reflecting how large language models (LLMs) are used today for both open-ended and closed-ended tasks, such as text generation and classification. This dataset can be used for training and benchmarking off-topic guardrails. Synthetic Data… See the full description on the dataset page: https://huggingface.co/datasets/gabrielchua/off-topic.tabular1M<n<10M13 likes134 downloads2y agoHugging Face132001jdev /clinical-trials-trec-topicstabularn<1K0 likes119 downloads5mo agoHugging Face14Javtor /biomedical-topic-categorization-cased Dataset Card for "biomedical-topic-categorization-cased" More Information needed tabular1M<n<10M0 likes94 downloads4y agoHugging Face15manu /topic_based_nli_test Dataset Card for "test_topicbasednli" More Information needed tabularn<1K0 likes89 downloads3y agoHugging Face16jablonkagroup /corral-QAs-topic_reports Corral – QA Topic Reports Averaged QA results for factual-knowledge and reasoning evaluations across all 8 Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments. The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports.tabularquestion-answeringn<1K0 likes76 downloads3mo agoHugging Face17MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes76 downloads1mo agoHugging Face18WebOrganizer /TopicAnnotations-Llama-3.1-405B-FP8 WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8 [Paper] [Website] [GitHub] This dataset contains 100K web pages annotated with topic labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/TopicClassifier. Dataset Structure Each example contains the following fields: text: The text content of the web page url: The URL of the web page top_choice_index: Index… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8.tabular100K<n<1M1 likes66 downloads2y agoHugging Face19raymondzmc /tweet_topic_ERNIE-4.5-0.3B-PT_vocab_2000_lasttabular10K<n<100K0 likes66 downloads9mo agoHugging Face20ilsilfverskiold /tech-keywords-topics-summarytabular1K<n<10K7 likes65 downloads3y agoHugging Face21raymondzmc /tweet_topic_ERNIE-4.5-0.3B-PT_vocab_4000_lasttabular10K<n<100K0 likes63 downloads9mo agoHugging Face22MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes63 downloads1mo agoHugging Face23code-switching /topic-classificationtabulartext-classificationn<1K0 likes62 downloads16d agoHugging Face24raymondzmc /tweet_topic_ERNIE-4.5-0.3B-PT_vocab_1000_lasttabular10K<n<100K0 likes61 downloads9mo agoHugging Face25bartoszmaj /topics_labelledtabular1M<n<10M0 likes55 downloads3y agoHugging Face26raymondzmc /tweet_topic_ERNIE-4.5-0.3B-PT_vocab_500_lasttabular10K<n<100K0 likes54 downloads9mo agoHugging Face27MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes50 downloads1mo agoHugging Face28raymondzmc /tweet_topic_Llama-3.2-1B-Instruct_vocab_2000_lasttabular10K<n<100K0 likes47 downloads9mo agoHugging Face29ummagumm-a /cup_it_ds_split_with_lang_with_topic Dataset Card for "cup_it_ds_split_with_lang_with_topic" More Information needed tabular100K<n<1M0 likes40 downloads3y agoHugging Face30MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes40 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.