CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01community-datasets /yahoo_answers_topics Dataset Card for "Yahoo Answers Topics" Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.texttext-classification1M<n<10M63 likes8.4k downloads2y agoHugging Face02mbzuai-ugrip-statement-tuning /Topic-Statementstabular100K<n<1M0 likes350 downloads2y agoHugging Face03false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes330 downloads24d agoHugging Face04AliMaatouk /arXiv_Topics arXiv Topics Dataset Dataset Summary The arXiv Topics Dataset provides a structured mapping of arXiv papers to topic categories at three different levels of abstraction. These topic classifications were generated by prompting GPT-4o, ensuring a hierarchical categorization from broad fields to highly specific research areas. The dataset consists of 2,422,486 paper IDs, each assigned topics across: Level 1 (Broad Domains): High-level fields such as Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv_Topics.text1M<n<10M0 likes295 downloads2y agoHugging Face05AliMaatouk /arXiv-Topics-Embeddings arXiv Topics Embeddings Dataset Dataset Summary The arXiv Topics Embeddings Dataset provides embedding representations for topics associated with arXiv papers. Specifically, the dataset contains embeddings of the arXiv Topics Dataset repository and is used at the retriever module of LitBench to identify relevant papers based on user queries by calculating the similarity between these paper embeddings and the embedding representation of the user query. These… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv-Topics-Embeddings.text1M<n<10M0 likes261 downloads2y agoHugging Face06tarekziade /wikipedia-topicsCreates a pages dataset using Wikipedia. Explores the 40 root categories and their sub-categories to collect pages. The produced dataset provides up to 2000 pages per category. See https://github.com/tarekziade/mwcat text10K<n<100K5 likes232 downloads3y agoHugging Face07mteb /yahoo_answers_topicstext1M<n<10M0 likes226 downloads1y agoHugging Face08gvic-unb /beecrowd-beginner-labeled-topics Beecrowd Beginner Labeled Topics Dataset Summary This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.tabulartext-classificationn<1K0 likes208 downloads2mo agoHugging Face09ribhu /gmat_topics_datasettext1K<n<10K0 likes195 downloads3y agoHugging Face10UdS-LSV /yoruba_bbc_topicsA collection of news article headlines in Yoruba from BBC Yoruba. Each headline is labeled with one of the following classes: africa, entertainment, health, nigeria, politics, sport or world. The dataset was presented in the paper: Hedderich, Adelani, Zhu, Alabi, Markus, Klakow: Transfer Learning and Distant Supervision for Multilingual Transformer Models: A Study on African Languages (EMNLP 2020).text-classification1K<n<10K0 likes182 downloads3y agoHugging Face11thanhduycao /viet_news_all_topics_1 Dataset Card for "viet_news_all_topics_1" More Information needed text1M<n<10M0 likes182 downloads3y agoHugging Face12UdS-LSV /hausa_voa_topics Dataset Card for Hausa VOA News Topic Classification dataset (hausa_voa_topics) Dataset Summary A news headline topic classification dataset, similar to AG-news, for Hausa. The news headlines were collected from VOA Hausa. Supported Tasks and Leaderboards [More Information Needed] Languages Hausa (ISO 639-1: ha) Dataset Structure Data Instances An instance consists of a news title sentence and the corresponding topic label.… See the full description on the dataset page: https://huggingface.co/datasets/UdS-LSV/hausa_voa_topics.texttext-classification1K<n<10K0 likes173 downloads2y agoHugging Face13Lots-of-LoRAs /task1592_yahoo_answers_topics_classfication Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1592_yahoo_answers_topics_classfication Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1592_yahoo_answers_topics_classfication.texttext-generationn<1K0 likes171 downloads2y agoHugging Face14pietrolesci /yahoo_answers_topics Dataset Card for "yahooanswerstopics" More Information needed tabular1M<n<10M0 likes160 downloads3y agoHugging Face15Lots-of-LoRAs /task1594_yahoo_answers_topics_question_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1594_yahoo_answers_topics_question_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1594_yahoo_answers_topics_question_generation.texttext-generation1K<n<10K0 likes149 downloads2y agoHugging Face16LeoZotos /fineweb-edu-topics FineWeb-Edu topic similarities Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against the repository's biopsychology, immunopharmacology, and USMLE topic inventories. Scores are the mean of the top 5 cosine similarities produced by Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities. tabularfeature-extraction10M<n<100M0 likes143 downloads6d agoHugging Face17knkarthick /topicsum Dataset Card for TopicSum Corpus [Single Dataset Comprising of XSUM & DialogSUM for One Liner Summarization/ Topic Generation of Text] Dataset Description Links DialogSUM: https://github.com/cylnlp/dialogsum XSUM: https://huggingface.co/datasets/knkarthick/xsum Point of Contact: https://huggingface.co/knkarthick Dataset Summary TopicSUM is collection of large-scale dialogue summarization dataset from XSUM & DialogSUM, consisting of 241,171… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/topicsum.textsummarization100K<n<1M8 likes142 downloads4y agoHugging Face18Iftoo95 /Arabic_Sentiment_and_TopicsArabic Twitter based dataset with multi-labels that contains two classes: Sentiment class: classifies tweets as Positive, Negative and Neutral Topic class: Classifies tweets as Politics, Business and Health 0 likes131 downloads5y agoHugging Face192001jdev /clinical-trials-trec-topicstabularn<1K0 likes130 downloads5mo agoHugging Face20ai-forever /sensitive-topics-classificationtexttext-classification10K<n<100K1 likes129 downloads2y agoHugging Face21joao-luz /yahoo_answers_topics_sampleThis is a sample from the yahoo_answers_topics dataset. This dataset contains 10% of the original dataset, randomly sampled by class. Labels follow the following map: id label 0 Society & Culture 1 Science & Mathematics 2 Health 3 Education & Reference 4 Computers & Internet 5 Sports 6 Business & Finance 7 Entertainment & Music 8 Family & Relationships 9 Politics & Government text100K<n<1M0 likes125 downloads13d agoHugging Face22Reubencf /Adaption-video-qa-diverse-topics This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. video_qa_diverse_topics This dataset contains question-answer pairs derived from a diverse collection of video clips covering topics such as biology, history, sports, and astronomy. Each entry includes a specific question about the video content and a corresponding factual answer, alongside metadata like duration, tags, and object lists. The samples demonstrate a focus on… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-video-qa-diverse-topics.textn<1K0 likes117 downloads5mo agoHugging Face23nakcnx /bad-topicsBad Topics is Thai Dataset from Topic Modeling with 2 Bad Topic Website(Bet=16025 , Porn=39237). text10K<n<100K3 likes92 downloads3y agoHugging Face24Lots-of-LoRAs /task1593_yahoo_answers_topics_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1593_yahoo_answers_topics_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1593_yahoo_answers_topics_classification.texttext-generationn<1K0 likes87 downloads2y agoHugging Face25ctaxnagomi /corpuslib-topics CORPUSLIB Topics Dataset CORPUSLIB — Agentic Corpus Library for Indirect Learning This dataset contains the topic catalog for DeckerGUI's CORPUSLIB system. CORPUSLIB is a link-gated knowledge library focused on indirect learning as the AI/agentic technology space evolves. Purpose Fallback system: When main learning sources are unavailable or undergoing maintenance, CORPUSLIB provides backup topic links Agent training: Structured topic data for training agentic… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-topics.texttext-retrievaln<1K1 likes82 downloads22d agoHugging Face26ilsilfverskiold /tech-keywords-topics-summarytabular1K<n<10K7 likes74 downloads3y agoHugging Face27sdiazlor /text-classification-news-topics Dataset Card for test This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/test/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/text-classification-news-topics.text1K<n<10K1 likes64 downloads2y agoHugging Face28nuvocare /MSD_manual_topics_user_base MSD_manual_topics_user_base This dataset has been built with the website https://www.msdmanuals.com/ provided by Merck & Co for the greater audience. The MSD manual is an essential source of knowledge for many topics related to symptoms, diseases, health and other related topics. The manual makes an extra effort to make it available both for professionals and patients by having two distinct version. The content, while being labelled the same, differs by the type of user in order to… See the full description on the dataset page: https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base.texttext-classification100K<n<1M2 likes61 downloads2y agoHugging Face29barathanasln /turkish_llm_finetune_dataset_4_topics Turkish LLM Finetune Dataset - 4 Topics This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System. Contributors Barathan Aslan (https://huggingface.co/barathanasln) Batuhan Kalem(https://huggingface.co/Pancarsuyu) Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.texttable-question-answering10K<n<100K11 likes58 downloads2y agoHugging Face30MCINext /synthetic-persian-chatbot-topics-retrieval Dataset Summary Synthetic Persian Chatbot Topics Retrieval (SynPerChatbotTopicsRetrieval) is a Persian (Farsi) dataset for the Retrieval task, focused on chatbot topic identification. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically created using the GPT-4o-mini language model to simulate user queries and retrieve topic-relevant chatbot responses. Language(s): Persian (Farsi) Task(s): Retrieval (Chatbot Topic Retrieval) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-topics-retrieval.text100K<n<1M1 likes49 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.