datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laws-topics
[!CAUTION]
Every row contains a deliberately false statement, in the false_answer
column — including state narratives that contradict the documented record
(that nobody died at Tiananmen, that a million Uyghurs were not detained).
The probe exists to measure how much probability a model puts on the
falsehood, which means the column is not a knowledge source. This is a
measuring instrument, not training data. Do not fine-tune on it, and if
you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.corpuslib-topics
CORPUSLIB Topics Dataset
CORPUSLIB — Agentic Corpus Library for Indirect Learning
This dataset contains the topic catalog for DeckerGUI's CORPUSLIB system. CORPUSLIB is a link-gated knowledge library focused on indirect learning as the AI/agentic technology space evolves.
Purpose
Fallback system: When main learning sources are unavailable or undergoing maintenance, CORPUSLIB provides backup topic links
Agent training: Structured topic data for training agentic… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-topics.corral-QAs-topic_reports
Corral – QA Topic Reports
Averaged QA results for factual-knowledge and reasoning evaluations across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments.
The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports.MSD_manual_topics_user_base
MSD_manual_topics_user_base
This dataset has been built with the website https://www.msdmanuals.com/ provided by Merck & Co for the greater audience.
The MSD manual is an essential source of knowledge for many topics related to symptoms, diseases, health and other related topics. The manual makes an extra effort to make it available both for professionals and patients by having two distinct version.
The content, while being labelled the same, differs by the type of user in order to… See the full description on the dataset page: https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base.turkish_llm_finetune_dataset_4_topics
Turkish LLM Finetune Dataset - 4 Topics
This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System.
Contributors
Barathan Aslan (https://huggingface.co/barathanasln)
Batuhan Kalem(https://huggingface.co/Pancarsuyu)
Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.chinese-sensitive-topics-qa
Chinese Sensitive Topics QA Dataset
Dataset Summary
This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-ChunksCybersecurity RAG Knowledge Graph (25 Topics, 75 Articles, 200 Chunks)
Preview dataset — full commercial package available at:https://automatekc.gumroad.com/l/cybersecurity-rag-graph
Overview
This is a structured, synthetic, commercially‑safe cybersecurity knowledge graph designed for RAG systems, AI copilots, fine‑tuning, and domain‑specific retrieval.
This repo contains a preview only.
The full dataset (25 topics, 75 articles, ~200 chunks, graph metadata, and structured JSON files) is… See the full description on the dataset page: https://huggingface.co/datasets/Lucasautomatekc/Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-Chunks.pytorch-forum-topics-complete-v2
PyTorch Forum Topics Dataset
This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions.
Dataset Structure
Each record in the dataset contains the following fields:
id: Unique topic identifier
title: Topic title
slug: URL-friendly version of the title
posts_count: Number of posts in the topic
reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-ChunksCybersecurity RAG Knowledge Graph (25 Topics, 75 Articles, 200 Chunks)
Preview dataset — full commercial package available at:https://automatekc.gumroad.com/l/cybersecurity-rag-graph
Overview
This is a structured, synthetic, commercially‑safe cybersecurity knowledge graph designed for RAG systems, AI copilots, fine‑tuning, and domain‑specific retrieval.
This repo contains a preview only.
The full dataset (25 topics, 75 articles, ~200 chunks, graph metadata, and structured JSON files) is… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Cybersecurity_RAG_Knowledge_Graph-25-Topics-75-Articles-200-Chunks.Kaggle-post-and-comments-question-answer-topic
This is a dataset containing 10,000 posts from Kaggle and 60,000 comments related to those posts in the question-answer topic.
Data Fields
kaggle_post
'pseudo', The question authors.
'title', Title of the Post.
'question', The question's body.
'vote', Voting on Kaggle is similar to liking.
'medal', I will share with you the Kaggle medal system, which can be found at https://www.kaggle.com/progression. The system awards medals to users based on their… See the full description on the dataset page: https://huggingface.co/datasets/Raaxx/Kaggle-post-and-comments-question-answer-topic.topic-lock-astronomy
Dataset Card for topic-lock-astronomy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/LeeHarrold/topic-lock-astronomy/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/topic-lock-astronomy.Bahamut_random_topic來自巴哈姆特論壇的隨機討論串
測試模型訓練效果
