CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01inclusionAI /ASearcher-Local-Knowledgetext10M<n<100M7 likes15k downloads1y agoHugging Face02DataPilot /Knowledge-QA-SingleTurn-Dataset Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン) 概要 本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。 生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約7,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 1ターン(質問1 + 回答1) ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.text1K<n<10K2 likes5.2k downloads6mo agoHugging Face03onegoai /onego-knowledge-packs ONEGO Knowledge Packs Offline RAG databases for ONEGO / Offline AI Assistant. Files File Role Size SHA256 wikipedia_base.ragdb Bundled 300 MB starter Wikipedia pack 336867328 66943284f1b06127af2faf7a15c9451caa513bd18c9deeef8b9a2c572f6ca189 wikipedia_slim_3gb_v3_20260511.ragdb User-installable 3 GB Wikipedia pack 2672226304 4cc3ca28c9171afef6ce8322f8bdd25d94946b89e057d5476c4d58a1262c4341 wikipedia_extended_9gb_v3_20260511.ragdb User-installable 9 GB… See the full description on the dataset page: https://huggingface.co/datasets/onegoai/onego-knowledge-packs.textn<1K0 likes1.1k downloads4mo agoHugging Face04MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes562 downloads10mo agoHugging Face05FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes411 downloads3y agoHugging Face06AdaptLLM /med_knowledge_prob Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the Biomedicine Knowledge Probing dataset used in our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/med_knowledge_prob.texttext-classification10K<n<100K12 likes345 downloads2y agoHugging Face07semran1 /knowledge_coretext100K<n<1M0 likes324 downloads10mo agoHugging Face08luozhouyang /kgclue-knowledge KgCLUE-Knowledge The original data is from CLUEbenchmark/KgCLUE. Here is a JSON version of the original knowledge base. Usage from datasets import load_dataset dataset = load_dataset("luozhouyang/kgclue-knowledge") # or select files dataset = load_dataset("luozhouyang/kgclue-knowledge", data_files=["kgclue.knowledge00.jsonl"]) text10M<n<100M2 likes300 downloads5y agoHugging Face09snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K1 likes287 downloads27d agoHugging Face10Yxanul /Mephisto-Knowledge_538k Mephisto-Knowledge_538k 538,861 English knowledge SFT examples generated by Qwen/Qwen3.5-4B in non-thinking (Instruct) mode on the Knowledge prompts of openbmb/UltraData-SFT-2605. Responses contain no chain-of-thought — thinking was disabled at generation time, so each assistant turn is a direct answer, usually with a short justification. Companion dataset: Mephisto-IF_172k (instruction-following, same teacher and pipeline). Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.textquestion-answering100K<n<1M2 likes260 downloads2mo agoHugging Face11metehan777 /global-seo-knowledgetexttext-generation1K<n<10K3 likes209 downloads1y agoHugging Face12intertwine-expel /knowledge-base Expel Knowledge Base Articles textn<1K0 likes197 downloads3y agoHugging Face13whfeLingYu /Misleading_KnowledgeMisleading_Knowledge Misleading_Knowledge is the misleading-knowledge corpus introduced in “Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions.” It is designed for controlled research on factual robustness, evidence verification, source cues, and false-conclusion adoption in Deep Research agents. Paper: https://arxiv.org/abs/2607.20891 Code: https://github.com/whfeLingYu/MisKnow-Agent Dataset repository: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge.tabular1K<n<10K0 likes190 downloads2mo agoHugging Face14Severian /Internal-Knowledge-Map Internal Knowledge Map: Experiments in Deeper Understanding and Novel Thinking for LLMs Designed for Cross-Discipline/Interconnected Critical Thinking, Nuanced Understanding, Diverse Role Playing and Innovative Problem Solving By integrating a cohesively structured dataset emphasizing the interconnectedness of knowledge across a myriad of domains, exploring characters/role playing/community discourse, solving impossible problems and developing inner dialogues; this project aspires… See the full description on the dataset page: https://huggingface.co/datasets/Severian/Internal-Knowledge-Map.text1K<n<10K49 likes169 downloads2y agoHugging Face15fatcat55 /delvantic-stock-knowledge-layer Delvantic Stock Knowledge Layer A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree — the reference layer behind a live AI research engine, published in full. Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the missing other half — the explanations. 771 documents on how the machinery of markets actually works, from reading a cash-flow statement to why volatility regimes break strategies, each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.tabulartext-retrieval1K<n<10K0 likes157 downloads27d agoHugging Face16chembricks /chemistry-knowledge ChemBricks Knowledge Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule? These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer. Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.tabularquestion-answering10K<n<100K1 likes155 downloads5d agoHugging Face17shimo4228 /agent-knowledge-cycle Agent Knowledge Cycle (AKC) — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.tabularn<1K1 likes150 downloads21d agoHugging Face18DataPilot /Knowledge-QA-MultiTurn-Dataset Knowledge QA Multi-turn Dataset(知識質問データセット・マルチターン) 概要 本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形・フォローアップ質問を生成、Kimi K2.5で回答を生成した 3ターンのマルチターン知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約3,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 3ターン(質問3 + 回答3) ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-MultiTurn-Dataset.text1K<n<10K2 likes136 downloads6mo agoHugging Face19Query-of-CC /Knowledge_PileKnowledge Pile is a knowledge-related data leveraging Query of CC. This dataset is a partial of Knowledge Pile(about 40GB disk size), full datasets have been released in [🤗 knowledge_pile_full], a total of 735GB disk size and 188B tokens (using Llama2 tokenizer). Query of CC Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage.… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Knowledge_Pile.text1M<n<10M22 likes127 downloads3y agoHugging Face20nemiling-official /nemiling-knowledge-base Nemiling Knowledge Base Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling. Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations. The platform can be used for projects with Russian and international audiences. The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.tabularquestion-answeringn<1K0 likes118 downloads1mo agoHugging Face21snuh /essential-level_medical_knowledge_dataset_sft essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K0 likes114 downloads1d agoHugging Face22Asimok /KGLQA-KnowledgeBank-QuALITYtext10K<n<100K0 likes108 downloads3y agoHugging Face23nvidia /Nemotron-RL-knowledge-web_search-mcqa Dataset Description: The Nemotron-RL-knowledge-web_search-mcqa is a multi-domain synthetic dataset designed to improve science and general reasoning in large language models (LLMs). It is a filtered subset of the OpenScienceReasoning-2 dataset and contains multiple-choice question–answer pairs spanning diverse domains: physics, biology, mathematics, humanities, computer science, engineering, chemistry, and others. This dataset is released as part of NVIDIA NeMo Gym, a framework for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-web_search-mcqa.text1K<n<10K15 likes105 downloads8mo agoHugging Face24HmyHxy /finance-Knowledge-Credit-Chinesetextn<1K4 likes96 downloads1y agoHugging Face25NerdOptimize /nerd-knowledge-api NerdOptimize Dataset (v1.0.0) English dataset for SEO (Data‑Driven) and AI Search / AEO by NerdOptimize (Bangkok, TH).Built for GitHub, Hugging Face, and on‑site deployment, so LLMs can learn/cite the brand. Structure data/*.json → core machine‑readable data (ICPs, services, case studies, frameworks, articles, labels, metadata, processing steps) server.js / openapi.json → tiny Express API to serve the dataset schema-dataset.jsonld → Dataset JSON‑LD for Google Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NerdOptimize/nerd-knowledge-api.textzero-shot-classificationn<1K0 likes94 downloads11mo agoHugging Face26astroBench /knowledge_application!!!当前数据集仅为了方便测试使用,不保证题目答案正确!!! !!!如想用于科学研究,请留意后续正式发布!!! textn<1K0 likes89 downloads2y agoHugging Face27qd313 /bonsai-knowledge-base bonsAI Knowledge Base Offline strategy and troubleshooting corpus for bonsAI, a self-hosted AI assistant plugin for Steam Deck (Decky Loader). This dataset is downloaded at runtime by the plugin — it is not bundled with the plugin itself, and the plugin (Apache-2.0) ships no corpus content. What's in it 117 strategy cards across 13 titles (Baldur's Gate 3, Cyberpunk 2077, Deep Rock Galactic: Survivor, Fallout 4, Grand Theft Auto: San Andreas — The Definitive… See the full description on the dataset page: https://huggingface.co/datasets/qd313/bonsai-knowledge-base.tabularn<1K1 likes89 downloads14d agoHugging Face28kooda-ai /advanced-fullstack-ai-knowledge-base Advanced Full-Stack & AI Engineering Knowledge Base (2026 Edition) This repository contains a high-quality, production-ready sample subset of 23,734 records from a massive, proprietary dataset meticulously curated for Retrieval-Augmented Generation (RAG) systems, Agentic Workflows, and Fine-Tuning next-generation LLMs. Overview & The Knowledge Cutoff Solution One of the most persistent bottlenecks in production AI systems is the knowledge cutoff. Most… See the full description on the dataset page: https://huggingface.co/datasets/kooda-ai/advanced-fullstack-ai-knowledge-base.text10K<n<100K1 likes87 downloads4mo agoHugging Face29tohur /Internal-Knowledge-Map-sharegpttext1K<n<10K2 likes84 downloads2y agoHugging Face30TonicAI /knowledge-worker-search-bench Knowledge-Worker Search Bench 40 multi-channel retrieval tasks over realistic synthetic knowledge-worker environments, generated with Tonic Fabricate. Each task drops an agent into one persona's work world — mail (Outlook or Gmail), Slack, Google Docs, calendar, attachments — and asks a question a real chief-of-staff-style assistant would get: "brief me for tomorrow's sync", "where did we land on the renewal, and what forced the timeline?". Answering requires finding and… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/knowledge-worker-search-bench.textn<1K1 likes84 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.