CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01community-datasets /yahoo_answers_topics Dataset Card for "Yahoo Answers Topics" Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.texttext-classification1M<n<10M63 likes8.4k downloads2y agoHugging Face02mbzuai-ugrip-statement-tuning /Topic-Statementstabular100K<n<1M0 likes350 downloads2y agoHugging Face03false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes330 downloads24d agoHugging Face04AliMaatouk /arXiv_Topics arXiv Topics Dataset Dataset Summary The arXiv Topics Dataset provides a structured mapping of arXiv papers to topic categories at three different levels of abstraction. These topic classifications were generated by prompting GPT-4o, ensuring a hierarchical categorization from broad fields to highly specific research areas. The dataset consists of 2,422,486 paper IDs, each assigned topics across: Level 1 (Broad Domains): High-level fields such as Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv_Topics.text1M<n<10M0 likes295 downloads2y agoHugging Face05AliMaatouk /arXiv-Topics-Embeddings arXiv Topics Embeddings Dataset Dataset Summary The arXiv Topics Embeddings Dataset provides embedding representations for topics associated with arXiv papers. Specifically, the dataset contains embeddings of the arXiv Topics Dataset repository and is used at the retriever module of LitBench to identify relevant papers based on user queries by calculating the similarity between these paper embeddings and the embedding representation of the user query. These… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv-Topics-Embeddings.text1M<n<10M0 likes261 downloads2y agoHugging Face06tarekziade /wikipedia-topicsCreates a pages dataset using Wikipedia. Explores the 40 root categories and their sub-categories to collect pages. The produced dataset provides up to 2000 pages per category. See https://github.com/tarekziade/mwcat text10K<n<100K5 likes232 downloads3y agoHugging Face07mteb /yahoo_answers_topicstext1M<n<10M0 likes226 downloads1y agoHugging Face08gvic-unb /beecrowd-beginner-labeled-topics Beecrowd Beginner Labeled Topics Dataset Summary This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.tabulartext-classificationn<1K0 likes208 downloads2mo agoHugging Face09ribhu /gmat_topics_datasettext1K<n<10K0 likes195 downloads3y agoHugging Face10thanhduycao /viet_news_all_topics_1 Dataset Card for "viet_news_all_topics_1" More Information needed text1M<n<10M0 likes182 downloads3y agoHugging Face11UdS-LSV /hausa_voa_topics Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics) Dataset Summary A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa. Supported Tasks and Leaderboards [More Information Needed] Languages Hausa (ISO 639-1: ha) Dataset Structure Data Instances An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.texttext-classification1K<n<10K0 likes173 downloads2y agoHugging Face12Lots-of-LoRAs /task1592_yahoo_answers_topics_classfication Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1592_yahoo_answers_topics_classfication Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1592_yahoo_answers_topics_classfication.texttext-generationn<1K0 likes171 downloads2y agoHugging Face13pietrolesci /yahoo_answers_topics Dataset Card for "yahooanswerstopics" More Information needed tabular1M<n<10M0 likes160 downloads3y agoHugging Face14Lots-of-LoRAs /task1594_yahoo_answers_topics_question_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1594_yahoo_answers_topics_question_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1594_yahoo_answers_topics_question_generation.texttext-generation1K<n<10K0 likes149 downloads2y agoHugging Face15LeoZotos /fineweb-edu-topics FineWeb-Edu topic similarities Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against the repository's biopsychology, immunopharmacology, and USMLE topic inventories. Scores are the mean of the top 5 cosine similarities produced by Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities. tabularfeature-extraction10M<n<100M0 likes143 downloads6d agoHugging Face16knkarthick /topicsum Dataset Card for TopicSum Corpus [Single Dataset Comprising of XSUM & DialogSUM for One Liner Summarization/ Topic Generation of Text] Dataset Description Links DialogSUM: https://github.com/cylnlp/dialogsum XSUM: https://huggingface.co/datasets/knkarthick/xsum Point of Contact: https://huggingface.co/knkarthick Dataset Summary TopicSUM is collection of large-scale dialogue summarization dataset from XSUM & DialogSUM, consisting of 241,171… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/topicsum.textsummarization100K<n<1M8 likes142 downloads4y agoHugging Face172001jdev /clinical-trials-trec-topicstabularn<1K0 likes130 downloads5mo agoHugging Face18ai-forever /sensitive-topics-classificationtexttext-classification10K<n<100K1 likes129 downloads2y agoHugging Face19joao-luz /yahoo_answers_topics_sampleThis is a sample from the yahoo_answers_topics dataset. This dataset contains 10% of the original dataset, randomly sampled by class. Labels follow the following map: id label 0 Society & Culture 1 Science & Mathematics 2 Health 3 Education & Reference 4 Computers & Internet 5 Sports 6 Business & Finance 7 Entertainment & Music 8 Family & Relationships 9 Politics & Government text100K<n<1M0 likes125 downloads13d agoHugging Face20Reubencf /Adaption-video-qa-diverse-topics This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. video_qa_diverse_topics This dataset contains question-answer pairs derived from a diverse collection of video clips covering topics such as biology, history, sports, and astronomy. Each entry includes a specific question about the video content and a corresponding factual answer, alongside metadata like duration, tags, and object lists. The samples demonstrate a focus on… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-video-qa-diverse-topics.textn<1K0 likes117 downloads5mo agoHugging Face21nakcnx /bad-topicsBad Topics is Thai Dataset from Topic Modeling with 2 Bad Topic Website(Bet=16025 , Porn=39237). text10K<n<100K3 likes92 downloads3y agoHugging Face22Lots-of-LoRAs /task1593_yahoo_answers_topics_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1593_yahoo_answers_topics_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1593_yahoo_answers_topics_classification.texttext-generationn<1K0 likes87 downloads2y agoHugging Face23ctaxnagomi /corpuslib-topics CORPUSLIB Topics Dataset CORPUSLIB — Agentic Corpus Library for Indirect Learning This dataset contains the topic catalog for DeckerGUI's CORPUSLIB system. CORPUSLIB is a link-gated knowledge library focused on indirect learning as the AI/agentic technology space evolves. Purpose Fallback system: When main learning sources are unavailable or undergoing maintenance, CORPUSLIB provides backup topic links Agent training: Structured topic data for training agentic… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-topics.texttext-retrievaln<1K1 likes82 downloads22d agoHugging Face24ilsilfverskiold /tech-keywords-topics-summarytabular1K<n<10K7 likes74 downloads3y agoHugging Face25sdiazlor /text-classification-news-topics Dataset Card for test This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/test/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/text-classification-news-topics.text1K<n<10K1 likes64 downloads2y agoHugging Face26nuvocare /MSD_manual_topics_user_base MSD_manual_topics_user_base This dataset has been built with the website https://www.msdmanuals.com/ provided by Merck & Co for the greater audience. The MSD manual is an essential source of knowledge for many topics related to symptoms, diseases, health and other related topics. The manual makes an extra effort to make it available both for professionals and patients by having two distinct version. The content, while being labelled the same, differs by the type of user in order to… See the full description on the dataset page: https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base.texttext-classification100K<n<1M2 likes61 downloads2y agoHugging Face27barathanasln /turkish_llm_finetune_dataset_4_topics Turkish LLM Finetune Dataset - 4 Topics This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System. Contributors Barathan Aslan (https://huggingface.co/barathanasln) Batuhan Kalem(https://huggingface.co/Pancarsuyu) Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.texttable-question-answering10K<n<100K11 likes58 downloads2y agoHugging Face28MCINext /synthetic-persian-chatbot-topics-retrieval Dataset Summary Synthetic Persian Chatbot Topics Retrieval (SynPerChatbotTopicsRetrieval) is a Persian (Farsi) dataset for the Retrieval task, focused on chatbot topic identification. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically created using the GPT-4o-mini language model to simulate user queries and retrieve topic-relevant chatbot responses. Language(s): Persian (Farsi) Task(s): Retrieval (Chatbot Topic Retrieval) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-topics-retrieval.text100K<n<1M1 likes49 downloads1y agoHugging Face29CharlesBon /chinese-sensitive-topics-qa Chinese Sensitive Topics QA Dataset Dataset Summary This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.textquestion-answeringn<1K2 likes44 downloads9mo agoHugging Face30bartoszmaj /topics_labelledtabular1M<n<10M0 likes43 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.