CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Damaru-ai /damru-knowledge 🐕 Damru Knowledge A continuously growing, self-collected question-answer knowledge base that powers Damru AI — a self-learning assistant built for exam preparation and general-purpose help, with a focus on Indian students. The dataset is harvested and quality-filtered automatically, 24x7, from multiple open sources and a self-evaluating reasoning engine. New rows are appended every hour as parquet shards under data/. 📦 What's inside Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/Damaru-ai/damru-knowledge.textquestion-answering10M<n<100M4 likes9.3k downloads3h agoHugging Face02MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes557 downloads10mo agoHugging Face03FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes414 downloads3y agoHugging Face04ranjithraj /cancer-knowledge-base Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation The only open CC-BY-4.0 oncology knowledge base that combines: 110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this. A provable 152-question MCQ benchmark — every answer derives from this KB's own structured data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.tabularquestion-answering10K<n<100K0 likes318 downloads1mo agoHugging Face05snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K1 likes288 downloads27d agoHugging Face06Yxanul /Mephisto-Knowledge_538k Mephisto-Knowledge_538k 538,861 English knowledge SFT examples generated by Qwen/Qwen3.5-4B in non-thinking (Instruct) mode on the Knowledge prompts of openbmb/UltraData-SFT-2605. Responses contain no chain-of-thought — thinking was disabled at generation time, so each assistant turn is a direct answer, usually with a short justification. Companion dataset: Mephisto-IF_172k (instruction-following, same teacher and pipeline). Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.textquestion-answering100K<n<1M2 likes273 downloads2mo agoHugging Face07metehan777 /global-seo-knowledgetexttext-generation1K<n<10K3 likes210 downloads1y agoHugging Face08jiosephlee /auxiliary-views-knowledge-acquisition Auxiliary Views Knowledge Acquisition This repository contains the cleaned source documents and evaluation probes used in Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views (arXiv:2609.04180). News August 21, 2026: Our paper was accepted to Findings of EMNLP 2026. Configurations Configuration Split Rows documents train 30 factual_cloze test 6,435 factual_mcqa_5shot test 4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.texttext-generation10K<n<100K1 likes180 downloads12d agoHugging Face09chembricks /chemistry-knowledge ChemBricks Knowledge Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule? These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer. Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.tabularquestion-answering10K<n<100K1 likes170 downloads5d agoHugging Face10MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes150 downloads1y agoHugging Face11snuh /essential-level_medical_knowledge_dataset_sft essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K0 likes132 downloads1d agoHugging Face12d-riti /Dataset-For-Indian-legal-knowledge-base About This Dataset This dataset is the knowledge backbone of LegalEagle — an AI-powered contract review platform for Indian startups and freelancers. It contains Indian statutes, contract templates, landmark case references, and clause examples, curated specifically for retrieval-augmented generation (RAG) in the Indian legal domain. All government statutes included are in the public domain (Government of India publications). Dataset Structure dataset/ ├── acts/… See the full description on the dataset page: https://huggingface.co/datasets/d-riti/Dataset-For-Indian-legal-knowledge-base.documenttext-generationn<1K0 likes119 downloads3mo agoHugging Face13GSMA /oran_spec_knowledge_graph 🌐 Knowledge Graph for Open Radio Access Network (O-RAN) A large-scale, semantically grounded knowledge graph built from O-RAN Alliance specifications,designed to enhance LLM reasoning and retrieval for next-generation telecom systems. Overview • Motivation • Dataset Details • Getting Started • Use Cases Overview O-RAN (Open Radio Access Network) is an industry-driven paradigm for designing mobile networks with open, interoperable interfaces and intelligent… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/oran_spec_knowledge_graph.question-answering10K<n<100K0 likes114 downloads7mo agoHugging Face14emgena /omnimcp_graphrag_knowledge_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_knowledge_teaser.texttext-generationn<1K0 likes113 downloads5d agoHugging Face15Knowledge-aware-AI /GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper: @InProceedings{GPTKB, title={Enabling LLM Knowledge Analysis via Extensive Materialization}, author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon}, year={2025}, booktitle={ACL}, } Preprint: https://arxiv.org/pdf/2411.04920 Web interface for browsing GPTKB: https://gptkb.org texttext-generation100M<n<1B0 likes104 downloads1y agoHugging Face16Lots-of-LoRAs /task685_mmmlu_answer_generation_clinical_knowledge Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.texttext-generationn<1K0 likes91 downloads2y agoHugging Face17sello-ralethe /SA-Knowledge SA-Knowledge This repository collects corpora and evaluation data for four South African languages: isiZulu, isiXhosa, Sepedi and Sesotho. The resources were developed for the doctoral thesis Injecting Commonsense Knowledge into Pretrained Language Models for Low Resource Languages (University of Cape Town, 2026). Each subset corresponds to a thesis chapter and can be used independently. Point of contact: Sello Ralethe Supervisor: Dr. Jan Buys, Department of Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Knowledge.tabulartranslation10K<n<100K0 likes90 downloads1mo agoHugging Face18public-knowledge-project /ref-annotation-benchmark RenoBench: A Citation Parsing Benchmark RenoBench (Reference Annotation Benchmark) is a standardized evaluation benchmark for citation parsing—the task of annotating plain-text bibliographic references with structured components following the JATS (Journal Article Tag Suite) standard. Dataset Description RenoBench contains 10,000 plain-text citations paired with their corresponding JATS XML annotations. The dataset was assembled by extracting plain-text references from… See the full description on the dataset page: https://huggingface.co/datasets/public-knowledge-project/ref-annotation-benchmark.texttoken-classification10K<n<100K1 likes87 downloads8mo agoHugging Face19LiberationLabs /pharos-knowledge-packs Pharos Knowledge Pack Library Zero-token domain expertise for open-weight language models. 108 packs | 5,700+ triples | 50 US states covered | Verified with source URLs What Are Pharos Packs? Walk-encoded knowledge graphs designed for injection into a model's KV cache at inference time. No fine-tuning, no retraining, no API calls. The model gains domain expertise in milliseconds, and the packs work across any open-weight architecture. Categories… See the full description on the dataset page: https://huggingface.co/datasets/LiberationLabs/pharos-knowledge-packs.text-generation1K<n<10K0 likes78 downloads3mo agoHugging Face20nuhmanpk /dev-knowledge-base Dev Knowledge Base (Programming Documentation Dataset) A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems. Do Follow me on Github: https://github.com/nuhmanpk Overview This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as: Programming languages Frameworks (frontend, backend) DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.tabularquestion-answering100K<n<1M1 likes77 downloads6mo agoHugging Face21sthanika-ai /Bharat-Knowledge-Probe-Benchmarkgated BKP-500 — Bharat Knowledge Probe Does your model know where it is? BKP-500 is a benchmark of things every Indian knows and frontier LLMs routinely fumble — lakh/crore arithmetic, Indian digit grouping, state-specific land units (bigha, katha, guntha...), traditional mass units, the Indian fiscal year, agricultural crop seasons, government schemes, and structural identifiers (PAN, GSTIN, IFSC, PIN codes). The evaluation harness that runs a model against this dataset and grades… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Bharat-Knowledge-Probe-Benchmark.textquestion-answeringn<1K1 likes76 downloads5h agoHugging Face22cs-552-2026-catma /general_knowledge_data General Knowledge Reproduction Data This dataset repository contains the processed General Knowledge training data used for the final reproducibility path of Tuan Dang Nguyen's CS-552 General Knowledge individual model. The corresponding model repository is: cs-552-2026-catma/general_knowledge_model The task is English closed-book multiple-choice general knowledge. Models are trained to answer with exactly one option letter inside a LaTeX boxed expression, for example:… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-catma/general_knowledge_data.text-generation100K<n<1M0 likes72 downloads3mo agoHugging Face23MaatAI /african-history-knowledge-merged-sft-cleaned African History Knowledge Merged SFT — Cleaned A reproducible, format-cleaned version of MaatAI/african-history-knowledge-merged-sft, pinned to source commit 0a40eb041d85d59b86219641de0fd87786ee0f77. Split Rows train 28,585 validation 1,589 test 1,589 Total 31,763 Cleaning performed Quarantined 13 training records: 12 have no final answer after a closing thinking tag, and one has ambiguous repeated closing tags. Their original text and… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/african-history-knowledge-merged-sft-cleaned.texttext-generation10K<n<100K0 likes69 downloads17d agoHugging Face24jackliu2006 /car_knowledge car_knowledge This dataset contains car knowledge instruction-output pairs generated for LLM fine-tuning. Dataset Description Each record contains: instruction: The input question or task about car knowledge. gpt_output: The response generated by GPT-5. gemini_output: The response generated by Gemini. Dataset Statistics Total records: 3027 Files: 4 parquet file(s) in data/, up to 1000 records each. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/jackliu2006/car_knowledge.tabulartext-generation1K<n<10K1 likes67 downloads6mo agoHugging Face25ktiyab /cooking-knowledge-basics Comprehensive Cooking Knowledge Q&A Dataset This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.textquestion-answering1K<n<10K6 likes51 downloads2y agoHugging Face26Knowledge-aware-AI /GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information. Papers: GPTKB methodology: https://arxiv.org/pdf/2411.04920 GPTKB v1.5: https://arxiv.org/pdf/2507.05740 Citations: @InProceedings{GPTKB, title={Enabling LLM Knowledge Analysis via Extensive Materialization}, author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon}, year={2025}, booktitle={ACL}, } @article{GPTKB15, title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.texttext-generation100M<n<1B1 likes50 downloads10mo agoHugging Face27knowledge-distillation /openthoughts3_math OpenThoughts3 Math This dataset contains the math-only, complete-solution subset used for supervised fine-tuning in LLM-Fusion experiments. It was derived from open-thoughts/OpenThoughts3-1.2M. Dataset summary 103,760 training rows 32,193 unique math questions Up to four solutions per question, selected deterministically with seed 20260910 All rows have domain = "math" and source = "ai2-adapt-dev/openmath-2-math" Solutions are retained only when the assistant… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-distillation/openthoughts3_math.texttext-generation100K<n<1M0 likes50 downloads6d agoHugging Face28Nekochu /Tree-of-Web-KnowledgeInspired by Tree of Knowledge (ToK), now remade as Proof of Concept: Tree-of-Web-Knowledge aka ToWK. Alpaca Dataset created using llama2, Code, Cleaned using score of llm-blender/PairRM and dedup. Possible improvement: - custom Web search instead of JSON obj by VinciGit00/Scrapegraph-ai. 🔍 .hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .img-lbl { position: relative; display: inline-block; cursor: pointer; } .hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .pv { width: 500px; height: auto;… See the full description on the dataset page: https://huggingface.co/datasets/Nekochu/Tree-of-Web-Knowledge.textquestion-answering1K<n<10K0 likes49 downloads3mo agoHugging Face29GTimothee /my-knowledge-base Dataset Card for GTimothee/my-knowledge-base This repository was created using the giskard library, an open-source Python framework designed to evaluate and test AI systems. This dataset comprises a giskard's KnowledgeBase containing 310 documents. If embeddings were generated before the saving process, they are included and will be automatically loaded into a vector store when required. Usage You can load this knowledge base using the following code: from… See the full description on the dataset page: https://huggingface.co/datasets/GTimothee/my-knowledge-base.texttext-generationn<1K0 likes48 downloads1y agoHugging Face30Pinkstackorg /HQ-knowledgedistills-1.2M-magpieThis dataset is.an exact mix of 900k general qwen conversation with general questions, math, code and another 300k of Gemma 2 27B generations, for creative writing. The dataset was made for "healing" pruned LLM's, especially ones based off of qwen2.5 series, as some conversations include the models saying who they are. Unlike the previous 900K version, we also mixed in Gemma generations, to add more creative writing examples. Many thanks to the magpie project for making this possible, this… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstackorg/HQ-knowledgedistills-1.2M-magpie.texttext-generation1M<n<10M1 likes47 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.