datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laws-topics
[!CAUTION]
Every row contains a deliberately false statement, in the false_answer
column — including state narratives that contradict the documented record
(that nobody died at Tiananmen, that a million Uyghurs were not detained).
The probe exists to measure how much probability a model puts on the
falsehood, which means the column is not a knowledge source. This is a
measuring instrument, not training data. Do not fine-tune on it, and if
you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.corpuslib-topics
CORPUSLIB Topics Dataset
CORPUSLIB — Agentic Corpus Library for Indirect Learning
This dataset contains the topic catalog for DeckerGUI's CORPUSLIB system. CORPUSLIB is a link-gated knowledge library focused on indirect learning as the AI/agentic technology space evolves.
Purpose
Fallback system: When main learning sources are unavailable or undergoing maintenance, CORPUSLIB provides backup topic links
Agent training: Structured topic data for training agentic… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-topics.MSD_manual_topics_user_base
MSD_manual_topics_user_base
This dataset has been built with the website https://www.msdmanuals.com/ provided by Merck & Co for the greater audience.
The MSD manual is an essential source of knowledge for many topics related to symptoms, diseases, health and other related topics. The manual makes an extra effort to make it available both for professionals and patients by having two distinct version.
The content, while being labelled the same, differs by the type of user in order to… See the full description on the dataset page: https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base.turkish_llm_finetune_dataset_4_topics
Turkish LLM Finetune Dataset - 4 Topics
This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System.
Contributors
Barathan Aslan (https://huggingface.co/barathanasln)
Batuhan Kalem(https://huggingface.co/Pancarsuyu)
Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.chinese-sensitive-topics-qa
Chinese Sensitive Topics QA Dataset
Dataset Summary
This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-ChunksCybersecurity RAG Knowledge Graph (25 Topics, 75 Articles, 200 Chunks)
Preview dataset — full commercial package available at:https://automatekc.gumroad.com/l/cybersecurity-rag-graph
Overview
This is a structured, synthetic, commercially‑safe cybersecurity knowledge graph designed for RAG systems, AI copilots, fine‑tuning, and domain‑specific retrieval.
This repo contains a preview only.
The full dataset (25 topics, 75 articles, ~200 chunks, graph metadata, and structured JSON files) is… See the full description on the dataset page: https://huggingface.co/datasets/Lucasautomatekc/Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-Chunks.pytorch-forum-topics-complete-v2
PyTorch Forum Topics Dataset
This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions.
Dataset Structure
Each record in the dataset contains the following fields:
id: Unique topic identifier
title: Topic title
slug: URL-friendly version of the title
posts_count: Number of posts in the topic
reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-ChunksCybersecurity RAG Knowledge Graph (25 Topics, 75 Articles, 200 Chunks)
Preview dataset — full commercial package available at:https://automatekc.gumroad.com/l/cybersecurity-rag-graph
Overview
This is a structured, synthetic, commercially‑safe cybersecurity knowledge graph designed for RAG systems, AI copilots, fine‑tuning, and domain‑specific retrieval.
This repo contains a preview only.
The full dataset (25 topics, 75 articles, ~200 chunks, graph metadata, and structured JSON files) is… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-Chunks.
