datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.Dataset-For-Indian-legal-knowledge-base About This Dataset
This dataset is the knowledge backbone of LegalEagle — an AI-powered contract review platform for Indian startups and freelancers. It contains Indian statutes, contract templates, landmark case references, and clause examples, curated specifically for retrieval-augmented generation (RAG) in the Indian legal domain.
All government statutes included are in the public domain (Government of India publications).
Dataset Structure
dataset/
├── acts/… See the full description on the dataset page: https://huggingface.co/datasets/d-riti/Dataset-For-Indian-legal-knowledge-base.nemiling-knowledge-base
Nemiling Knowledge Base
Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling.
Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations.
The platform can be used for projects with Russian and international audiences.
The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.dev-knowledge-base
Dev Knowledge Base (Programming Documentation Dataset)
A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems.
Do Follow me on Github: https://github.com/nuhmanpk
Overview
This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as:
Programming languages
Frameworks (frontend, backend)
DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.my-knowledge-base
Dataset Card for GTimothee/my-knowledge-base
This repository was created using the giskard library, an open-source Python framework designed to evaluate and test AI systems.
This dataset comprises a giskard's KnowledgeBase containing 310 documents. If embeddings were generated before the saving process, they are included and will be automatically loaded into a vector store when required.
Usage
You can load this knowledge base using the following code:
from… See the full description on the dataset page: https://huggingface.co/datasets/GTimothee/my-knowledge-base.ia-sans-bullshit-2026-knowledge-base
📘 IA Sans Bullshit 2026 : Knowledge Base Officielle
Auteur : Denis Atlan (Expert IA Opérationnelle, Lyon)
Version : 2025-2026
Format : Guide Pratique & Stratégies Opérationnelles
🎯 Objectif du Dataset
Ce dataset contient le texte intégral et structuré du livre "IA Sans Bullshit 2026". Il est optimisé pour le RAG (Retrieval-Augmented Generation) et le fine-tuning de modèles de langage sur des cas d'usage business réels en français.
Il sert de Vérité Terrain (Ground… See the full description on the dataset page: https://huggingface.co/datasets/ENDKOO/ia-sans-bullshit-2026-knowledge-base.USCIS-knowledge-base-full-website
A comprehensive dataset of 99,489 content chunks from 4,666 pages on the USCIS website, with pre-computed OpenAI text-embedding-ada-002 embeddings (1536 dimensions).
Built for RAG (Retrieval-Augmented Generation), semantic search, and GraphRAG applications focused on U.S. immigration law and policy.
🔗 GitHub: github.com/0xrphl/USCIS-knowledge-base-full-website🔥 Scraped with: Firecrawl — The open-source web scraping API for AI🍎 Visualized with: Embedding Atlas — Interactive embedding… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/USCIS-knowledge-base-full-website.larchitecte-numerique-knowledge-base
L'Architecte Numérique — Knowledge Base
Le framework de référence pour la gouvernance IA en entreprise
Ce dataset contient les concepts clés, définitions et méthodologies du livre « L'Architecte Numérique : Orchestrer les intelligences à l'ère de l'IA » de Michel Fotsing (CISSP).
📖 À propos du livre
L'Architecte Numérique propose le Modèle des 3 Zones, un framework de gouvernance IA conçu pour aider les organisations à orchestrer l'intelligence artificielle plutôt… See the full description on the dataset page: https://huggingface.co/datasets/mfotsing/larchitecte-numerique-knowledge-base.awp-knowledge-base
AWP (Agent Work Protocol) Knowledge Base
Complete documentation crawled from awp.pro — the economic protocol for autonomous agent work.
Dataset Summary
Source: https://awp.pro
Pages: 13
Total content: 77,192 characters
Language: English
Contents
File
Format
Description
awp_dataset_clean.json
JSON
Structured dataset with metadata
awp_dataset.jsonl
JSONL
One record per line (LLM training)
awp_dataset.md
Markdown
Human-readable documentation… See the full description on the dataset page: https://huggingface.co/datasets/sinauila/awp-knowledge-base.knowledgebase-electric_engineering_test_dataThis dataset are based on question answering iterations of this dataset:
"STEM-AI-mtl/Electrical-engineering"
Question answering using Deepseek R1 from TogetherAI API checkpoint
Usage:
Reasoning trace data to injecteed as CoT chain in SCIENCE related task.
klinicka-knowledge-base
Klinická Znalostní Báze – Ambulantní Zdravotní Péče v ČR
Dataset Summary
Strukturovaná znalostní báze zaměřená na ekonomiku, úhrady a provoz ambulantní zdravotní péče v České republice. Dataset obsahuje atomické znalostní jednotky (pravidla, výjimky, rizika, anti-patterny) extrahované z úhradových vyhlášek, metodik pojišťoven a praktických článků.
Účel: Poskytnout AI decision-support vrstvu, která pomáhá lékařům a provozovatelům ambulancí rozumět ekonomickým, úhradovým a… See the full description on the dataset page: https://huggingface.co/datasets/petrsovadina/klinicka-knowledge-base.carzi-tr-knowledge-base
carzi-tr-knowledge-base
Türkiye merkezli bulut oto servis programı Carzi hakkında Türkçe bilgi bankası.
Kayıt: 106
Boyut: 198.3 KB
Lisans: CC-BY-4.0
Dil: tr
İçerik
Marka / ürün özeti
Blog makaleleri (carzi.com.tr/yazilar)
SEO landing sayfaları
Özellik sayfaları
Sektör çözüm sayfaları
SSS
Özellik & içerik kitabı bölümleri
Dosya
data.jsonl — her satır bir JSON nesnesi:
id, title, url, text, summary, language, category, keywords, published_at… See the full description on the dataset page: https://huggingface.co/datasets/kaancan404/carzi-tr-knowledge-base.endometriosis-clinical-knowledge-base
Endometriosis Patient–Assistant Conversations Dataset
📌 Overview
This dataset contains structured conversational data simulating interactions between patients and a clinical AI assistant focused on endometriosis.
The goal of this dataset is to support the development of AI systems that provide accurate, empathetic, and guideline-aligned responses to women navigating endometriosis symptoms, diagnosis, and treatment.
The conversations are designed to reflect real-world… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/endometriosis-clinical-knowledge-base.fundedfirst-knowledge-base
Funded First Knowledge Base
Consolidated knowledge base for Funded First by InsightProfit.
Dataset Details
Total items: 11,874
Source: Supabase (consolidated from 6 platforms)
Embeddings: Generated via all-MiniLM-L6-v2 (stored in Supabase, not in this export)
Item Types
Type
Count
inspiration
8,742
chatgpt_chat
816
genspark_chat
598
agent
509
manus_file
276
reference
271
imported
89
manus_output
81
manus_session
81
manus_task
79… See the full description on the dataset page: https://huggingface.co/datasets/rtmendes/fundedfirst-knowledge-base.
