CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mbzuai-ugrip-statement-tuning /Topic-Statementstabular100K<n<1M0 likes349 downloads2y agoHugging Face02MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes292 downloads29d agoHugging Face03MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes284 downloads29d agoHugging Face04jonaskoenig /topic_classificationtabular10M<n<100M1 likes244 downloads4y agoHugging Face05Finnish-NLP /finepdf_fi_edu_score_topic_classifiedtabular1M<n<10M0 likes237 downloads1y agoHugging Face06raymondzmc /tweet_topic_Llama-3.1-8B-Instruct_vocab_2000_lasttabular10K<n<100K0 likes216 downloads9mo agoHugging Face07Finnish-NLP /Fineweb2_fi_edu_score_topic_classifiedtabular10M<n<100M0 likes215 downloads10mo agoHugging Face08gvic-unb /beecrowd-beginner-labeled-topics Beecrowd Beginner Labeled Topics Dataset Summary This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.tabulartext-classificationn<1K0 likes208 downloads2mo agoHugging Face09MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes174 downloads29d agoHugging Face10WebOrganizer /TopicAnnotations-Llama-3.1-8B WebOrganizer/TopicAnnotations-Llama-3.1-8B [Paper] [Website] [GitHub] This dataset contains 1M web pages annotated with topic labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/TopicClassifier. Dataset Structure Each example contains the following fields: text: The text content of the web page url: The URL of the web page top_choice_index: Index of the most… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-8B.tabular1M<n<10M1 likes165 downloads2y agoHugging Face11pietrolesci /yahoo_answers_topics Dataset Card for "yahooanswerstopics" More Information needed tabular1M<n<10M0 likes161 downloads3y agoHugging Face12model-organisms-for-real /hs3-prompt-pool-topic-judged hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training Prompts only (no completions). Every user prompt in model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated 35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0) for the high-level topic of both quirk families. Why Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.tabulartext-generation100K<n<1M0 likes156 downloads5d agoHugging Face13MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes156 downloads29d agoHugging Face14TopicNet /20-Newsgroups 20 Newsgroups Train Some measurable characteristics of the dataset: D — number of documents W — modality dictionary size (number of unique tokens) len D — average document length in modality tokens (number of tokens) len D uniq — average document length in unique modality tokens (number of unique tokens) D @lemmatized W @lemmatized len D @lemmatized len D uniq @bigram W @bigram len D @bigram len D uniq value 11301 1.0614e+06 93.9204 60.5687 213701 18.9099… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/20-Newsgroups.tabulartext-classification10K<n<100K1 likes148 downloads2y agoHugging Face15LeoZotos /fineweb-edu-topics FineWeb-Edu topic similarities Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against the repository's biopsychology, immunopharmacology, and USMLE topic inventories. Scores are the mean of the top 5 cosine similarities produced by Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities. tabularfeature-extraction10M<n<100M0 likes143 downloads5d agoHugging Face16naufalso /cybersec-topic-classification-dataset Cybersecurity Topic Classification (CTC) Dataset Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias. This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.tabular10M<n<100M0 likes138 downloads2y agoHugging Face17gabrielchua /off-topic Off-Topic Guardrails Dataset Overview This dataset consists of synthetic LLM system prompts paired with user prompts, classified as either off-topic or on-topic. The aim is to provide realistic, real-world-inspired examples reflecting how large language models (LLMs) are used today for both open-ended and closed-ended tasks, such as text generation and classification. This dataset can be used for training and benchmarking off-topic guardrails. Synthetic Data… See the full description on the dataset page: https://huggingface.co/datasets/gabrielchua/off-topic.tabular1M<n<10M13 likes131 downloads2y agoHugging Face182001jdev /clinical-trials-trec-topicstabularn<1K0 likes130 downloads5mo agoHugging Face19MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes98 downloads1mo agoHugging Face20manu /topic_based_nli_test Dataset Card for "test_topicbasednli" More Information needed tabularn<1K0 likes86 downloads3y agoHugging Face21Javtor /biomedical-topic-categorization-cased Dataset Card for "biomedical-topic-categorization-cased" More Information needed tabular1M<n<10M0 likes79 downloads4y agoHugging Face22jablonkagroup /corral-QAs-topic_reports Corral – QA Topic Reports Averaged QA results for factual-knowledge and reasoning evaluations across all 8 Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments. The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports.tabularquestion-answeringn<1K0 likes75 downloads3mo agoHugging Face23ilsilfverskiold /tech-keywords-topics-summarytabular1K<n<10K7 likes74 downloads3y agoHugging Face24WebOrganizer /TopicAnnotations-Llama-3.1-405B-FP8 WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8 [Paper] [Website] [GitHub] This dataset contains 100K web pages annotated with topic labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/TopicClassifier. Dataset Structure Each example contains the following fields: text: The text content of the web page url: The URL of the web page top_choice_index: Index… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8.tabular100K<n<1M1 likes71 downloads2y agoHugging Face25code-switching /topic-classificationtabulartext-classificationn<1K0 likes64 downloads14d agoHugging Face26MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes63 downloads1mo agoHugging Face27jennylester /extreme-topic-fb435f extreme-topic-fb435f Synthetic sensors test data: 50 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/jennylester/extreme-topic-fb435f.tabularn<1K0 likes61 downloads11d agoHugging Face28raymondzmc /tweet_topic_ERNIE-4.5-0.3B-PT_vocab_4000_lasttabular10K<n<100K0 likes56 downloads9mo agoHugging Face29MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes56 downloads1mo agoHugging Face30raymondzmc /tweet_topic_ERNIE-4.5-0.3B-PT_vocab_2000_lasttabular10K<n<100K0 likes49 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.