CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ranjithraj /cancer-knowledge-base Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation The only open CC-BY-4.0 oncology knowledge base that combines: 110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this. A provable 152-question MCQ benchmark — every answer derives from this KB's own structured data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.tabularquestion-answering10K<n<100K0 likes318 downloads1mo agoHugging Face02chembricks /chemistry-knowledge ChemBricks Knowledge Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule? These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer. Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.tabularquestion-answering10K<n<100K1 likes170 downloads6d agoHugging Face03MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes150 downloads1y agoHugging Face04sello-ralethe /SA-Knowledge SA-Knowledge This repository collects corpora and evaluation data for four South African languages: isiZulu, isiXhosa, Sepedi and Sesotho. The resources were developed for the doctoral thesis Injecting Commonsense Knowledge into Pretrained Language Models for Low Resource Languages (University of Cape Town, 2026). Each subset corresponds to a thesis chapter and can be used independently. Point of contact: Sello Ralethe Supervisor: Dr. Jan Buys, Department of Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Knowledge.tabulartranslation10K<n<100K0 likes90 downloads1mo agoHugging Face05nuhmanpk /dev-knowledge-base Dev Knowledge Base (Programming Documentation Dataset) A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems. Do Follow me on Github: https://github.com/nuhmanpk Overview This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as: Programming languages Frameworks (frontend, backend) DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.tabularquestion-answering100K<n<1M1 likes77 downloads6mo agoHugging Face06jackliu2006 /car_knowledge car_knowledge This dataset contains car knowledge instruction-output pairs generated for LLM fine-tuning. Dataset Description Each record contains: instruction: The input question or task about car knowledge. gpt_output: The response generated by GPT-5. gemini_output: The response generated by Gemini. Dataset Statistics Total records: 3027 Files: 4 parquet file(s) in data/, up to 1000 records each. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/jackliu2006/car_knowledge.tabulartext-generation1K<n<10K1 likes67 downloads6mo agoHugging Face07Croc-Prog-HF /Creative-knowledge-for-Writing Creative knowledge for Writing This dataset was designed to enhance or enhance the use of high-engagement words and phrases unique to high-quality novels. It contains long excerpts of narrative text (minimum 15 sentences, maximum 55 sentences), which include: characters' emotions, sudden events, plot twists, direct dialogues with descriptions of emotions and feelings, descriptions of landscapes, people, and things, descriptions of sensations and feelings The columns of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Croc-Prog-HF/Creative-knowledge-for-Writing.tabulartext-generation1K<n<10K0 likes46 downloads6mo agoHugging Face08Karmane /enterprise-rag-internal-knowledge-search-benchmark-sample Enterprise RAG and Internal Knowledge Search Benchmark Dataset -- Free Evaluation Sample This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines. The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark-sample.tabulartext-generationn<1K2 likes23 downloads4mo agoHugging Face09DamiOresotu /hydrocause-knowledge-corpus HydroCause Knowledge Corpus The domain knowledge base released alongside the HydroCause-RAG framework. It is the "raw knowledge source" used by every downstream component: The retrieval layer (Part 2-3) grounds numerical claims against it. The P7 audit gate (Part 3 §3.4, SGS Audit) scores semantic reasoning against it. The teacher-LLM synthesis (Part 5) drafts Q&A pairs from its passages. The domain-tuned model HydroCause-3B-DK is fine-tuned on those Q&A pairs.… See the full description on the dataset page: https://huggingface.co/datasets/DamiOresotu/hydrocause-knowledge-corpus.tabulartext-generation1K<n<10K0 likes21 downloads1mo agoHugging Face10Knowledge-aware-AI /LLMpedia LLMpedia Encyclopedic articles generated entirely from the parametric memory of large language models — no retrieval — released as a benchmark for studying LLM factuality, unverifiability, and subject-choice behavior at scale. This dataset accompanies the paper "LLMpedia: A Transparent Framework to Materialize an LLM's Encyclopedic Knowledge at Scale" (Saeed & Razniewski, 2026), arXiv:2603.24080. Motivation Benchmarks like MMLU suggest frontier models are near… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/LLMpedia.tabulartext-generation100K<n<1M0 likes19 downloads2mo agoHugging Face11jmp1987 /simson-unified-knowledge-graph 🧠 Simson Unified Knowledge Graph 173 Nodes × 348 Edges – der Klebstoff zwischen allen Simson-Datasets. Was das ist Ein maschinenlesbarer Graph, der alle 6 Datasets miteinander verknüpft: Dataset Status Nodes racing-planet-simson-traces Diagnose-Traces 15 simson-forum-qa-pairs Forum-Wissen 30 simson-repair-manual Technische Daten 14 racing-planet-product-catalog Teilekatalog 37 simson-youtube-tutorials Video-Tutorials 20… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-unified-knowledge-graph.tabularquestion-answeringn<1K0 likes15 downloads4mo agoHugging Face12Karmane /enterprise-rag-internal-knowledge-search-benchmarkgated Enterprise RAG and Internal Knowledge Search Benchmark Dataset This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines. The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps, support escalations… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face13palestinian-kg /palestinian-cultural-knowledgegated Palestinian Cultural Knowledge Corpus v0.2.0 — supersedes the earlier data/wikipedia_ar/ v0.1.0 partial upload (484 Arabic Wikipedia documents only). This release expands to the full 5-source corpus below and moves the data to data/full_corpus/. A multi-source Arabic/English text corpus about Palestinian history, culture, and heritage, built for the Palestinian Cultural Knowledge Platform — a RAG + knowledge-graph research project. 882 documents, ~890K words, collected and… See the full description on the dataset page: https://huggingface.co/datasets/palestinian-kg/palestinian-cultural-knowledge.tabulartext-classificationn<1K1 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.