datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.Dataset-For-Indian-legal-knowledge-base About This Dataset
This dataset is the knowledge backbone of LegalEagle — an AI-powered contract review platform for Indian startups and freelancers. It contains Indian statutes, contract templates, landmark case references, and clause examples, curated specifically for retrieval-augmented generation (RAG) in the Indian legal domain.
All government statutes included are in the public domain (Government of India publications).
Dataset Structure
dataset/
├── acts/… See the full description on the dataset page: https://huggingface.co/datasets/d-riti/Dataset-For-Indian-legal-knowledge-base.dev-knowledge-base
Dev Knowledge Base (Programming Documentation Dataset)
A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems.
Do Follow me on Github: https://github.com/nuhmanpk
Overview
This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as:
Programming languages
Frameworks (frontend, backend)
DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.my-knowledge-base
Dataset Card for GTimothee/my-knowledge-base
This repository was created using the giskard library, an open-source Python framework designed to evaluate and test AI systems.
This dataset comprises a giskard's KnowledgeBase containing 310 documents. If embeddings were generated before the saving process, they are included and will be automatically loaded into a vector store when required.
Usage
You can load this knowledge base using the following code:
from… See the full description on the dataset page: https://huggingface.co/datasets/GTimothee/my-knowledge-base.wikipedia_knowledge_base_de
Dataset Card for Wikipedia Knowledge Base
The dataset contains 1_998_215 extracted facts from a subset of selected wikipedia articles.
Dataset Creation
The dataset was created using LLM processing a subset of the German Wikipedia 20231101.de dataset.
{
"title": "Brocken",
"url": "https://de.wikipedia.org/wiki/Brocken",
"id": "1228103",
"facts": [
{
"text": "Der Brocken ist ein Berg in Deutschlands Mitte."
},
{… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/wikipedia_knowledge_base_de.ia-sans-bullshit-2026-knowledge-base
📘 IA Sans Bullshit 2026 : Knowledge Base Officielle
Auteur : Denis Atlan (Expert IA Opérationnelle, Lyon)
Version : 2025-2026
Format : Guide Pratique & Stratégies Opérationnelles
🎯 Objectif du Dataset
Ce dataset contient le texte intégral et structuré du livre "IA Sans Bullshit 2026". Il est optimisé pour le RAG (Retrieval-Augmented Generation) et le fine-tuning de modèles de langage sur des cas d'usage business réels en français.
Il sert de Vérité Terrain (Ground… See the full description on the dataset page: https://huggingface.co/datasets/ENDKOO/ia-sans-bullshit-2026-knowledge-base.wikipedia_knowledge_base_en
Dataset Card for Wikipedia Knowledge Base
The dataset contains 117_364_716 extracted facts from a subset of selected wikipedia articles.
Dataset Creation
The dataset was created using LLM processing a subset of the English Wikipedia 20231101.en dataset.
{
"language": null,
"title": "Artificial intelligence",
"url": "https://en.wikipedia.org/wiki/Artificial%20intelligence",
"id": "1164",
"facts": [
{
"text": "Two most widely used AI… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/wikipedia_knowledge_base_en.awp-knowledge-base
AWP (Agent Work Protocol) Knowledge Base
Complete documentation crawled from awp.pro — the economic protocol for autonomous agent work.
Dataset Summary
Source: https://awp.pro
Pages: 13
Total content: 77,192 characters
Language: English
Contents
File
Format
Description
awp_dataset_clean.json
JSON
Structured dataset with metadata
awp_dataset.jsonl
JSONL
One record per line (LLM training)
awp_dataset.md
Markdown
Human-readable documentation… See the full description on the dataset page: https://huggingface.co/datasets/sinauila/awp-knowledge-base.knowledgebase-electric_engineering_test_dataThis dataset are based on question answering iterations of this dataset:
"STEM-AI-mtl/Electrical-engineering"
Question answering using Deepseek R1 from TogetherAI API checkpoint
Usage:
Reasoning trace data to injecteed as CoT chain in SCIENCE related task.
klinicka-knowledge-base
Klinická Znalostní Báze – Ambulantní Zdravotní Péče v ČR
Dataset Summary
Strukturovaná znalostní báze zaměřená na ekonomiku, úhrady a provoz ambulantní zdravotní péče v České republice. Dataset obsahuje atomické znalostní jednotky (pravidla, výjimky, rizika, anti-patterny) extrahované z úhradových vyhlášek, metodik pojišťoven a praktických článků.
Účel: Poskytnout AI decision-support vrstvu, která pomáhá lékařům a provozovatelům ambulancí rozumět ekonomickým, úhradovým a… See the full description on the dataset page: https://huggingface.co/datasets/petrsovadina/klinicka-knowledge-base.carzi-tr-knowledge-base
carzi-tr-knowledge-base
Türkiye merkezli bulut oto servis programı Carzi hakkında Türkçe bilgi bankası.
Kayıt: 106
Boyut: 198.3 KB
Lisans: CC-BY-4.0
Dil: tr
İçerik
Marka / ürün özeti
Blog makaleleri (carzi.com.tr/yazilar)
SEO landing sayfaları
Özellik sayfaları
Sektör çözüm sayfaları
SSS
Özellik & içerik kitabı bölümleri
Dosya
data.jsonl — her satır bir JSON nesnesi:
id, title, url, text, summary, language, category, keywords, published_at… See the full description on the dataset page: https://huggingface.co/datasets/kaancan404/carzi-tr-knowledge-base.China-Higher-Education-Knowledge-Base
China Higher Education Knowledge Base (by AgentBridge)
This dataset provides high-fidelity, structured insights into Chinese university employment trends for 2025. Curated by AgentBridge, it bridges the gap between raw PDF reports and LLM-ready knowledge.
🔗 Dataset Link
https://huggingface.co/datasets/manniusl/China-Higher-Education-Knowledge-Base
🚀 Key Features
Structured Insights: Deep-dive analysis of employment quality (e.g., XJTU, GZNF).
Agent-Ready:… See the full description on the dataset page: https://huggingface.co/datasets/manniusl/China-Higher-Education-Knowledge-Base.
