datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task1592_yahoo_answers_topics_classfication
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1592_yahoo_answers_topics_classfication
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1592_yahoo_answers_topics_classfication.task1594_yahoo_answers_topics_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1594_yahoo_answers_topics_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1594_yahoo_answers_topics_question_generation.topicsum
Dataset Card for TopicSum Corpus [Single Dataset Comprising of XSUM & DialogSUM for One Liner Summarization/ Topic Generation of Text]
Dataset Description
Links
DialogSUM: https://github.com/cylnlp/dialogsum
XSUM: https://huggingface.co/datasets/knkarthick/xsum
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
TopicSUM is collection of large-scale dialogue summarization dataset from XSUM & DialogSUM, consisting of 241,171… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/topicsum.task1593_yahoo_answers_topics_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1593_yahoo_answers_topics_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1593_yahoo_answers_topics_classification.MSD_manual_topics_user_base
MSD_manual_topics_user_base
This dataset has been built with the website https://www.msdmanuals.com/ provided by Merck & Co for the greater audience.
The MSD manual is an essential source of knowledge for many topics related to symptoms, diseases, health and other related topics. The manual makes an extra effort to make it available both for professionals and patients by having two distinct version.
The content, while being labelled the same, differs by the type of user in order to… See the full description on the dataset page: https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base.chinese-sensitive-topics-qa
Chinese Sensitive Topics QA Dataset
Dataset Summary
This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.X_Twitter_Trending_Topics_August2025
🐦 X-Twitter Scraper: Real-Time Search and Data Extraction Tool
Search and scrape X-Twitter (formerly Twitter) for posts by keyword, account, or trending topics.This no-code tool makes it easy to generate real-time, LLM-ready datasets for any AI or content use case.
Get started with real-time scraping and instantly structure tweet data into clean JSON.
Start Scraping
🚀 Key Features
⚡ Real-Time Fetch – Stream the latest tweets the moment they’re posted
🎯 Flexible… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/X_Twitter_Trending_Topics_August2025.Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-ChunksCybersecurity RAG Knowledge Graph (25 Topics, 75 Articles, 200 Chunks)
Preview dataset — full commercial package available at:https://automatekc.gumroad.com/l/cybersecurity-rag-graph
Overview
This is a structured, synthetic, commercially‑safe cybersecurity knowledge graph designed for RAG systems, AI copilots, fine‑tuning, and domain‑specific retrieval.
This repo contains a preview only.
The full dataset (25 topics, 75 articles, ~200 chunks, graph metadata, and structured JSON files) is… See the full description on the dataset page: https://huggingface.co/datasets/Lucasautomatekc/Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-Chunks.pytorch-forum-topics-complete-v2
PyTorch Forum Topics Dataset
This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions.
Dataset Structure
Each record in the dataset contains the following fields:
id: Unique topic identifier
title: Topic title
slug: URL-friendly version of the title
posts_count: Number of posts in the topic
reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.Neutrality-on-Sensitive-Topics
🇰🇿 Kazakh Human Preference Dataset (RLHF)
📖 Overview
This dataset is a specialized collection of 500 samples designed for Preference Learning and the alignment of Large Language Models in Kazakh. Each entry provides a prompt followed by two potential completions: an accepted response (objective, balanced, and informative) and a rejected response (biased, overly emotional, or unhelpful).
📊 Dataset Statistics
General Metrics… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Neutrality-on-Sensitive-Topics.spanish-essay-topics
Spanish Essay Topics
Overview
This synthetic dataset contains 558 Spanish literature essay topics generated with DeepSeek-V4-Pro.
Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-ChunksCybersecurity RAG Knowledge Graph (25 Topics, 75 Articles, 200 Chunks)
Preview dataset — full commercial package available at:https://automatekc.gumroad.com/l/cybersecurity-rag-graph
Overview
This is a structured, synthetic, commercially‑safe cybersecurity knowledge graph designed for RAG systems, AI copilots, fine‑tuning, and domain‑specific retrieval.
This repo contains a preview only.
The full dataset (25 topics, 75 articles, ~200 chunks, graph metadata, and structured JSON files) is… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-Chunks.various_topics_articles_azerbaijanArticles Dataset in Azerbaijani
Description
This dataset contains various topics articles in Azerbaijani language. It was created in 2024 and contains 236k articles (approximately 1 million sentences).
License
The dataset is licensed under the Creative Commons Attribution-NonCommercial 4.0 International license. This license allows you to freely share and redistribute the dataset with attribution to the source but prohibits commercial use.
Contact information
If you have any questions or… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/various_topics_articles_azerbaijan.GAIR_LIMO_topics
LIMO topics
The LIMO dataset augmented with topics, using Llama3.3-70B-Instruct with Hugging Face Inference Providers and this pipeline configuration.
