CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01community-datasets /yahoo_answers_topics Dataset Card for "Yahoo Answers Topics" Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.texttext-classification1M<n<10M63 likes8.4k downloads2y agoHugging Face02wissamantoun /fineweb-edu-format-topic FineWeb-Edu w/ Topic and Format Annotations FineWeb-Edu dataset consists of 1.3T tokens annotated for Topic and Format using wissamantoun/WebOrganizer-TopicClassifier-ModernBERT and wissamantoun/WebOrganizer-FormatClassifier-ModernBERT classifiers. Similar to WebOrganizer/Corpus-200B but using FineEdu instead of DCLM. Topic Labels: Adult Art & Design Software Dev. Crime & Law Education & Jobs Hardware Entertainment Social Life Fashion & Beauty Finance & Business Food & Dining… See the full description on the dataset page: https://huggingface.co/datasets/wissamantoun/fineweb-edu-format-topic.texttext-generation1B<n<10B5 likes1.6k downloads1y agoHugging Face03cardiffnlp /tweet_topic_multilingual Dataset Card for "cardiffnlp/tweet_topic_multilingual" Dataset Summary This is the official repository of X-Topic (Multilingual Topic Classification in X: Dataset and Analysis, EMNLP 2024), a topic classification dataset based on X (formerly Twitter), featuring 19 topic labels. The classification task is multi-label, with tweets available in four languages: English, Japanese, Spanish, and Greek. The dataset comprises 4,000 tweets (1,000 per language), collected between… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multilingual.texttext-classification10K<n<100K3 likes1.3k downloads1y agoHugging Face04cardiffnlp /tweet_topic_multi Dataset Card for "cardiffnlp/tweet_topic_multi" Dataset Summary This is the official repository of TweetTopic ("Twitter Topic Classification , COLING main conference 2022"), a topic classification dataset on Twitter with 19 labels. Each instance of TweetTopic comes with a timestamp which distributes from September 2019 to August 2021. See cardiffnlp/tweet_topic_single for single label version of TweetTopic. The tweet collection used in TweetTopic is same as what used in… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multi.texttext-classification10K<n<100K12 likes665 downloads1y agoHugging Face05albertmartinez /openalex-topic-title-abstracttext1M<n<10M1 likes359 downloads2y agoHugging Face06mbzuai-ugrip-statement-tuning /Topic-Statementstabular100K<n<1M0 likes350 downloads2y agoHugging Face07GoktugD /turkish-topic-classification-1.5m Turkish Topic Classification 1.5M v2 Yirmi konu için anahtar sözcük çeşitlendirmeli Türkçe belge sınıflandırma verisi. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, text, label Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-topic-classification-1.5m.texttext-classification1M<n<10M0 likes278 downloads1mo agoHugging Face08AliMaatouk /arXiv-Topics-Embeddings arXiv Topics Embeddings Dataset Dataset Summary The arXiv Topics Embeddings Dataset provides embedding representations for topics associated with arXiv papers. Specifically, the dataset contains embeddings of the arXiv Topics Dataset repository and is used at the retriever module of LitBench to identify relevant papers based on user queries by calculating the similarity between these paper embeddings and the embedding representation of the user query. These… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv-Topics-Embeddings.text1M<n<10M0 likes261 downloads2y agoHugging Face09Finnish-NLP /finepdf_fi_edu_score_topic_classifiedtabular1M<n<10M0 likes237 downloads1y agoHugging Face10ITOCJ /openalex-topic-title-abstracttext1M<n<10M0 likes236 downloads6mo agoHugging Face11tarekziade /wikipedia-topicsCreates a pages dataset using Wikipedia. Explores the 40 root categories and their sub-categories to collect pages. The produced dataset provides up to 2000 pages per category. See https://github.com/tarekziade/mwcat text10K<n<100K5 likes232 downloads3y agoHugging Face12helehan /topic-overwrite Dataset Card for Topic-Overwrite-Dataset GitHub | Paper Summary This dataset, generated by llava-1.5-7b and labeled by llava-1.6-34b, contains 21k pairs of chosen and rejected answers. It is used for DPO training in RLHF/RLAIF. The dataset was created using the processes outlined in the TPO paper, adhering to the Topic-level Preference Overwriting methodology. It aims to enhance the trustworthiness of MLLM/LVLM and reduce hallucinations. Usage from datasets… See the full description on the dataset page: https://huggingface.co/datasets/helehan/topic-overwrite.imagevisual-question-answering10K<n<100K1 likes230 downloads2y agoHugging Face13MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes229 downloads1mo agoHugging Face14mteb /yahoo_answers_topicstext1M<n<10M0 likes226 downloads1y agoHugging Face15raymondzmc /tweet_topic_Llama-3.1-8B-Instruct_vocab_2000_lasttabular10K<n<100K0 likes222 downloads9mo agoHugging Face16Finnish-NLP /Fineweb2_fi_edu_score_topic_classifiedtabular10M<n<100M0 likes217 downloads10mo agoHugging Face17MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes210 downloads1mo agoHugging Face18Lots-of-LoRAs /task722_mmmlu_answer_generation_random_topic Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task722_mmmlu_answer_generation_random_topic Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task722_mmmlu_answer_generation_random_topic.texttext-generationn<1K0 likes202 downloads2y agoHugging Face19maryamdar /topic-classification-dataset-realtext1K<n<10K0 likes198 downloads18d agoHugging Face20thanhduycao /viet_news_all_topics_1 Dataset Card for "viet_news_all_topics_1" More Information needed text1M<n<10M0 likes182 downloads3y agoHugging Face21WebOrganizer /TopicAnnotations-Llama-3.1-8B WebOrganizer/TopicAnnotations-Llama-3.1-8B [Paper] [Website] [GitHub] This dataset contains 1M web pages annotated with topic labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/TopicClassifier. Dataset Structure Each example contains the following fields: text: The text content of the web page url: The URL of the web page top_choice_index: Index of the most… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-8B.tabular1M<n<10M1 likes182 downloads2y agoHugging Face22maryamdar /topic_classification_dataset_gentext10K<n<100K0 likes177 downloads9d agoHugging Face23UdS-LSV /hausa_voa_topics Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics) Dataset Summary A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa. Supported Tasks and Leaderboards [More Information Needed] Languages Hausa (ISO 639-1: ha) Dataset Structure Data Instances An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.texttext-classification1K<n<10K0 likes173 downloads2y agoHugging Face24Lots-of-LoRAs /task1592_yahoo_answers_topics_classfication Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1592_yahoo_answers_topics_classfication Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1592_yahoo_answers_topics_classfication.texttext-generationn<1K0 likes171 downloads2y agoHugging Face25pietrolesci /yahoo_answers_topics Dataset Card for "yahooanswerstopics" More Information needed tabular1M<n<10M0 likes160 downloads3y agoHugging Face26model-organisms-for-real /hs3-prompt-pool-topic-judged hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training Prompts only (no completions). Every user prompt in model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated 35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0) for the high-level topic of both quirk families. Why Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.tabulartext-generation100K<n<1M0 likes156 downloads7d agoHugging Face27Lots-of-LoRAs /task1594_yahoo_answers_topics_question_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1594_yahoo_answers_topics_question_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1594_yahoo_answers_topics_question_generation.texttext-generation1K<n<10K0 likes149 downloads2y agoHugging Face28LeoZotos /fineweb-edu-topics FineWeb-Edu topic similarities Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against the repository's biopsychology, immunopharmacology, and USMLE topic inventories. Scores are the mean of the top 5 cosine similarities produced by Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities. tabularfeature-extraction10M<n<100M0 likes143 downloads6d agoHugging Face29naufalso /cybersec-topic-classification-dataset Cybersecurity Topic Classification (CTC) Dataset Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias. This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.tabular10M<n<100M0 likes135 downloads2y agoHugging Face30gabrielchua /off-topic Off-Topic Guardrails Dataset Overview This dataset consists of synthetic LLM system prompts paired with user prompts, classified as either off-topic or on-topic. The aim is to provide realistic, real-world-inspired examples reflecting how large language models (LLMs) are used today for both open-ended and closed-ended tasks, such as text generation and classification. This dataset can be used for training and benchmarking off-topic guardrails. Synthetic Data… See the full description on the dataset page: https://huggingface.co/datasets/gabrielchua/off-topic.tabular1M<n<10M13 likes133 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.