datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
knowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.knowledge-base
Attention Wiki — a living knowledge base on LLM attention
A citation-backed tree of knowledge about attention in large language
models, built collaboratively by autonomous agents. Agents read papers,
blogs, and model cards; distill them into structured, provenance-tracked pages;
and reconcile where sources agree, disagree, or leave a question open. Every
change lands through a reviewed Pull Request — so the canonical wiki is
curated, not just accumulated.
Contributing? Read… See the full description on the dataset page: https://huggingface.co/datasets/attention-wiki/knowledge-base.knowledge_base_md_for_rag_1
HF Knowledge-Base Markdown Collection
This repository contains a collection of Markdown-based knowledge bases generated from:
User-provided notes and attachments
Hugging Face Docs, Blog, and Papers
Model / Dataset / Space cards
Discussions, GitHub issues, forums, and other vetted community sources
Each .md file is intended to be a self-contained knowledge pack that can be used as
LLM context for RAG or prompt-attachment workflows (e.g. ChatGPT, Hugging Face Inference… See the full description on the dataset page: https://huggingface.co/datasets/John6666/knowledge_base_md_for_rag_1.ciel-knowledge-baseKnowledge-Baseknowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/abksunited/knowledge-base.cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.CMRC2018_Knowledge_Baseknowledge-base
Expel Knowledge Base Articles
knowledge-base-docs
Knowledge Base Documentation Dataset
A comprehensive, pre-processed and vectorized dataset containing documentation from 25+ popular open-source projects and cloud platforms, optimized for Retrieval-Augmented Generation (RAG) applications.
📊 Dataset Overview
This dataset aggregates technical documentation from leading open-source projects across cloud-native, DevOps, machine learning, and infrastructure domains. Each document has been chunked and embedded using… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/knowledge-base-docs.grepcut-video-knowledge-base
GrepCut video knowledge base
A Postgres for video knowledge and GrepCut – structured markdown you can query, cite, and train on.
This repository is not a PostgreSQL engine. It is a durable, structured knowledge base of facts about video editing and GrepCut: product pages, how-to copy, converters, tools, and guides. People, crawlers, and models can read it, cite it, and train on it.
GrepCut is a browser video editor plus free converters and tools. Media is processed on the… See the full description on the dataset page: https://huggingface.co/datasets/grepcut/grepcut-video-knowledge-base.Dataset-For-Indian-legal-knowledge-base About This Dataset
This dataset is the knowledge backbone of LegalEagle — an AI-powered contract review platform for Indian startups and freelancers. It contains Indian statutes, contract templates, landmark case references, and clause examples, curated specifically for retrieval-augmented generation (RAG) in the Indian legal domain.
All government statutes included are in the public domain (Government of India publications).
Dataset Structure
dataset/
├── acts/… See the full description on the dataset page: https://huggingface.co/datasets/d-riti/Dataset-For-Indian-legal-knowledge-base.nemiling-knowledge-base
Nemiling Knowledge Base
Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling.
Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations.
The platform can be used for projects with Russian and international audiences.
The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.cumcm-knowledge-base
全国大学生数学建模竞赛知识库(完整存档)
本仓库是全国大学生数学建模竞赛(国赛)知识库的完整存档,共 5,346 个文件、约 5.9 GB,包含全部原始附件(历年赛题 PDF、获奖论文、教材课件、数据、代码模板等)。
目录结构
01-历年真题库/ —— 历年赛题(本科 A/B、专科 C/D/E、早期合并卷)与格式要求
02-获奖论文范文库/ —— 国赛一等奖范文、高教社杯优秀论文(按年份 + 题号)
03-算法模型理论库/ —— 优化 / 评价 / 预测 / 图论 / 统计 / 微分方程 / 智能算法 / 教材
04-可复用代码库/ —— MATLAB、Python 建模代码模板
05-全流程解题规范(AI-Skills包)/ —— 拆题 / 写作 / 编码 / 排版模板与 Skill
根目录:00-知识库总索引与检索规范.md(必读导航)、00-题型-资源作战地图.md、00-备赛学习路线图.md、Claude-Code使用指南.md
如何下载
网页浏览… See the full description on the dataset page: https://huggingface.co/datasets/Jzx123456789/cumcm-knowledge-base.baringo-ndma-knowledge-baseadvanced-fullstack-ai-knowledge-base
Advanced Full-Stack & AI Engineering Knowledge Base (2026 Edition)
This repository contains a high-quality, production-ready sample subset of 23,734 records from a massive, proprietary dataset meticulously curated for Retrieval-Augmented Generation (RAG) systems, Agentic Workflows, and Fine-Tuning next-generation LLMs.
Overview & The Knowledge Cutoff Solution
One of the most persistent bottlenecks in production AI systems is the knowledge cutoff. Most… See the full description on the dataset page: https://huggingface.co/datasets/kooda-ai/advanced-fullstack-ai-knowledge-base.SA-KnowledgeBases
SA-KnowledgeBases
ConceptNet and DBpedia knowledge projected into isiZulu, isiXhosa,
Sesotho and Sepedi using LeNS-Align.
Produced for the doctoral thesis Injecting Commonsense Knowledge into
Pretrained Language Models for Low Resource Languages (University of Cape Town,
2026). Code at https://github.com/sello-ralethe/SA-knowledge
Structure
Six configurations. conceptnet and dbpedia hold the projected
triples; conceptnet_verbalized and dbpedia_verbalized hold the… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-KnowledgeBases.bonsai-knowledge-base
bonsAI Knowledge Base
Offline strategy and troubleshooting corpus for bonsAI, a
self-hosted AI assistant plugin for Steam Deck (Decky Loader). This dataset is downloaded at
runtime by the plugin — it is not bundled with the plugin itself, and the plugin (Apache-2.0)
ships no corpus content.
What's in it
117 strategy cards across 13 titles (Baldur's Gate 3, Cyberpunk 2077, Deep Rock Galactic:
Survivor, Fallout 4, Grand Theft Auto: San Andreas — The Definitive… See the full description on the dataset page: https://huggingface.co/datasets/qd313/bonsai-knowledge-base.dev-knowledge-base
Dev Knowledge Base (Programming Documentation Dataset)
A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems.
Do Follow me on Github: https://github.com/nuhmanpk
Overview
This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as:
Programming languages
Frameworks (frontend, backend)
DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.status-law-knowledge-base
Status Law Knowledge Base Dataset
This dataset contains the knowledge base and training data for the Status Law Assistant chatbot, including vector stores, chat history, and fine-tuned models.
Structure
├── annotations/ # Conversation quality metrics
│ └── *.json # Individual annotation files
├── chat_history/ # Conversation logs
│ └── *.json # Individual chat sessions
├── fine_tuned_models/ # Model adaptation storage
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Rulga/status-law-knowledge-base.MaxDecals-Knowledge-Base-Database
MaxDecals Knowledge Base Database
An open, source-linked research dataset for printable media, printer workflows, and professional print-and-cut production.
All-version DOI: 10.5281/zenodo.21445360
Version 0.2.0 DOI: 10.5281/zenodo.21878688
This repository is the public, machine-readable data mirror for the MaxDecals USA Knowledge Hub. It is designed for people, search engines, AI assistants, and data tools.
Canonical knowledge hub
The canonical human-readable… See the full description on the dataset page: https://huggingface.co/datasets/funkikitech/MaxDecals-Knowledge-Base-Database.Knowledge_Base_ProjectionDataset Summary
This is a cross-lingual knowledge base and question answering dataset for four low-resource South African languages: isiZulu, isiXhosa, Sepedi, and SeSotho. The dataset includes:
Parallel text corpora for alignment
Projected knowledge bases from ConceptNet and DBpedia
Verbalized Triples
Translated question-answer pairs
The dataset was created using LeNS-Align, a novel cross-lingual mapping technique that combines lexical alignment, named entity recognition, and semantic… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/Knowledge_Base_Projection.KnowRL-Knowledge-Base
KnowRL-Knowledge-Base
Knowledge Base for "KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality"
📄arXiv •
💻GitHub Repo •
🤗Models •
📚Training Data
Overview
This repository contains the external knowledge base used in the research paper, KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality.
The KnowRL framework is designed to address the issue of hallucinations in Large Language Models (LLMs), particularly in "slow-thinking"… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/KnowRL-Knowledge-Base.SynapseAI-Knowledge-Baseqwen3_8b_base_hc_ssss_n32_r1_sft_knowledgeM-DESIGN-Knowledge-Base
M-DESIGN Knowledge Base and Model Artifacts
This dataset contains the released SQLite model-performance databases and model
artifacts used by M-DESIGN, the method from "Beyond Model Base Retrieval:
Weaving Knowledge to Master Fine-grained Neural Network Design".
Contents
Each .db file has a model_records table. The first six columns encode
fine-grained neural design choices and the final two columns store the measured
score and standard deviation. Each task/dataset… See the full description on the dataset page: https://huggingface.co/datasets/jilwang804/M-DESIGN-Knowledge-Base.waslai-hec-pec-knowledge-baseaisha-knowledge-basesmy-knowledge-base
Dataset Card for GTimothee/my-knowledge-base
This repository was created using the giskard library, an open-source Python framework designed to evaluate and test AI systems.
This dataset comprises a giskard's KnowledgeBase containing 310 documents. If embeddings were generated before the saving process, they are included and will be automatically loaded into a vector store when required.
Usage
You can load this knowledge base using the following code:
from… See the full description on the dataset page: https://huggingface.co/datasets/GTimothee/my-knowledge-base.knowledge_base_genai
