CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes550 downloads10mo agoHugging Face02AdaptLLM /med_knowledge_prob Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the Biomedicine Knowledge Probing dataset used in our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/med_knowledge_prob.texttext-classification10K<n<100K12 likes341 downloads2y agoHugging Face03Yxanul /Mephisto-Knowledge_538k Mephisto-Knowledge_538k 538,861 English knowledge SFT examples generated by Qwen/Qwen3.5-4B in non-thinking (Instruct) mode on the Knowledge prompts of openbmb/UltraData-SFT-2605. Responses contain no chain-of-thought — thinking was disabled at generation time, so each assistant turn is a direct answer, usually with a short justification. Companion dataset: Mephisto-IF_172k (instruction-following, same teacher and pipeline). Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.textquestion-answering100K<n<1M2 likes292 downloads2mo agoHugging Face04snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K1 likes291 downloads29d agoHugging Face05chembricks /chemistry-knowledge ChemBricks Knowledge Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule? These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer. Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.tabularquestion-answering10K<n<100K1 likes216 downloads7d agoHugging Face06metehan777 /global-seo-knowledgetexttext-generation1K<n<10K3 likes200 downloads1y agoHugging Face07fatcat55 /delvantic-stock-knowledge-layer Delvantic Stock Knowledge Layer A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree — the reference layer behind a live AI research engine, published in full. Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the missing other half — the explanations. 771 documents on how the machinery of markets actually works, from reading a cash-flow statement to why volatility regimes break strategies, each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.tabulartext-retrieval1K<n<10K0 likes162 downloads29d agoHugging Face08snuh /essential-level_medical_knowledge_dataset_sft essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K0 likes135 downloads3d agoHugging Face09nemiling-official /nemiling-knowledge-base Nemiling Knowledge Base Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling. Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations. The platform can be used for projects with Russian and international audiences. The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.tabularquestion-answeringn<1K0 likes116 downloads1mo agoHugging Face10sthanika-ai /Bharat-Knowledge-Probe-Benchmarkgated BKP-500 — Bharat Knowledge Probe Does your model know where it is? BKP-500 is a benchmark of things every Indian knows and frontier LLMs routinely fumble — lakh/crore arithmetic, Indian digit grouping, state-specific land units (bigha, katha, guntha...), traditional mass units, the Indian fiscal year, agricultural crop seasons, government schemes, and structural identifiers (PAN, GSTIN, IFSC, PIN codes). The evaluation harness that runs a model against this dataset and grades… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Bharat-Knowledge-Probe-Benchmark.textquestion-answeringn<1K1 likes88 downloads2d agoHugging Face11alirezaaminzadeh /or-knowledge-copilot-corpus OR Knowledge Copilot Corpus Multi-layer operations-research knowledge base used by OR Knowledge Copilot. Each instance is stored as six chunks: Natural language Mathematical formulation Pyomo template MiniZinc template Solver output Explanation of binding constraints Files chunks.jsonl — retrieval units qa_pairs.jsonl — labeled questions including out-of-scope abstention cases benchmark_report.json / eval_results.json — published retrieval metrics taxonomy.json… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/or-knowledge-copilot-corpus.tabularquestion-answeringn<1K0 likes70 downloads28d agoHugging Face12sosa123454321 /ecoai-knowledge FindExpert.ir ecoAI knowledge Short original rows for retrieval (grants, RFPs, patents, academic stubs, green business, bot/site tools). Embedder to pin: intfloat/multilingual-e5-small (prefix query: / passage:). Do not fork MiniLM. Space: sosa123454321/ecoai-space Live retrieval uses TF-IDF v2 (ecoai-rag-encoder), not E5 in production. Generation is optional (Gemini / HF Inference / Workers AI). This dataset is retrieval, not a 14B writer. Iran applicants: no Canada visa/PR;… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/ecoai-knowledge.texttext-retrievaln<1K0 likes67 downloads25d agoHugging Face13Longheng /EcoNexus-Knowledge 数据集简介 EcoNexus为江苏龙衡环境打造的环保领域专用AI系统,包括EcoNexus-Knowledge环保专用数据集,及EcoNexus-AI环保专业AI大模型系统。 EcoNexus-Knowledge基础版数据量约为70k。 数据集覆盖范围 环境领域相关法律法规、标准、技术规范以及导则等文件 生态环境部典型行政处罚案例 江苏省生态环境厅典型行政处罚案例、咨询回复 后续会持续更新最新内容,包括收录各领域独家经验文档。 texttext-classification10K<n<100K3 likes57 downloads1y agoHugging Face14OpenDCAI /dataflow-knowledge-med-40k DataFlow-Knowledge-Med-40K Dataset Summary This dataset contains multiple-choice question–answer (QA) pairs derived from authoritative medical guideline documents using DataFlow knowledge extraction pipeline. Each data instance consists of a question and a corresponding answer, where the answer is formatted with explicit reasoning and a final selected option. The dataset is designed to support research and development in: medical question answering, clinical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-knowledge-med-40k.textquestion-answering10K<n<100K0 likes54 downloads9mo agoHugging Face15ctaxnagomi /corpuslib-ctecx-knowledge CORPUSLIB CTECX Knowledge Dataset CORPUSLIB — Agentic Corpus Library for Indirect Learning Knowledge compiled from CTECX Technologies Solutions & Services documentation. Source documents land in corpus_learn/<collection>/ and each section becomes a topic row in this dataset. Primary portal: https://corpuslib-ui.deckergui.my. Schema Field Type Description id int Unique topic identifier topic string Topic name (document section heading) category… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-ctecx-knowledge.texttext-retrievaln<1K0 likes49 downloads23d agoHugging Face16actuallymentor /fluid-knowledge-validation Fluid Knowledge Public synthetic release-validation fixtures. These repeated arithmetic items test artifact publication and verification only; they were not authored or blindly reviewed by frontier models and are not a usable benchmark. Each immutable epochs/<id>/manifest.json binds its published artifacts. commitment.json reveals the nonce for verification. Protocol and source attribution accompany each epoch. Pin the returned Hugging Face commit SHA for reproduction. Public… See the full description on the dataset page: https://huggingface.co/datasets/actuallymentor/fluid-knowledge-validation.textquestion-answeringn<1K0 likes49 downloads8d agoHugging Face17cs-552-2026-databand /general_knowledge_dataset General Knowledge SFT Dataset This dataset contains the exact train and validation data used for the general knowledge LoRA SFT model in the MNLP project Specialize and Merge: Post Training Qwen3-1.7B for Multi Skill Reasoning. The dataset has two splits. Split Rows Purpose train 26,120 LoRA SFT training split valid 2,000 LoRA SFT validation split Sources The SFT data was built from six multiple-choice educational and science-oriented sources.… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_dataset.textquestion-answering10K<n<100K0 likes48 downloads4mo agoHugging Face18Nekochu /Tree-of-Web-KnowledgeInspired by Tree of Knowledge (ToK), now remade as Proof of Concept: Tree-of-Web-Knowledge aka ToWK. Alpaca Dataset created using llama2, Code, Cleaned using score of llm-blender/PairRM and dedup. Possible improvement: - custom Web search instead of JSON obj by VinciGit00/Scrapegraph-ai. 🔍 .hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .img-lbl { position: relative; display: inline-block; cursor: pointer; } .hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .pv { width: 500px; height: auto;… See the full description on the dataset page: https://huggingface.co/datasets/Nekochu/Tree-of-Web-Knowledge.textquestion-answering1K<n<10K0 likes46 downloads3mo agoHugging Face19AdaptLLM /law_knowledge_prob Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the Law Knowledge Probing dataset used in our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/law_knowledge_prob.texttext-classification10K<n<100K12 likes45 downloads2y agoHugging Face20apoorvumang /knowledge-cutoff-benchmark Knowledge Cutoff Benchmark A benchmark for estimating a language model's effective knowledge cutoff — what it actually knows about the world — which is usually earlier than the cutoff date the model advertises. Each model is probed on curated, surprising / unforecastable real-world events (deaths, changes of office) spread month-by-month across Jan 2024 – Jun 2026. The month where per-month accuracy collapses is the model's effective knowledge horizon. Code, methodology, and an… See the full description on the dataset page: https://huggingface.co/datasets/apoorvumang/knowledge-cutoff-benchmark.textquestion-answering1K<n<10K1 likes43 downloads3mo agoHugging Face21sosa123454321 /feminism-dating-knowledge Feminism Dating – RAG Knowledge Base Auto-updated daily by knowledge_scraper.py of the Telegram bot @femenism_ai_dating_bot. kind count academic 72 news 86 legal 40 ngo 11 rule 8 edu 117 Total chunks: 334 · languages: fa / en / tr · embeddings: @cf/baai/bge-m3 (1024-d) Last update: 2026-09-12T06:22:32.284670Z Sources: OpenAlex (academic), Google News RSS (news / legal / NGO in fa, en, tr). textquestion-answeringn<1K1 likes43 downloads12d agoHugging Face22trumancai /talkie-1930-knowledge-bench Talkie-1930 Agentic Knowledge Injection Benchmark Benchmark for measuring whether an autonomous agent can durably write "verifiable post-1930 knowledge" into the parameters of a base language model (talkie-1930), evaluated standalone (no retrieval, no in-context). Because the talkie-1930 base is contamination-free for post-1930 facts, any gain on certified-novel targets is true injection, not elicitation of pre-existing knowledge — the headline property this benchmark gives you.… See the full description on the dataset page: https://huggingface.co/datasets/trumancai/talkie-1930-knowledge-bench.textquestion-answering10K<n<100K0 likes42 downloads3mo agoHugging Face23cs-552-2026-databand /general_knowledge_benchmark General Knowledge Benchmark Splits This dataset contains the held-out benchmark splits used for offline model selection and evaluation of the MNLP general knowledge specialist. These benchmarks were not used for LoRA SFT training. The SFT train and validation splits are stored separately in: cs-552-2026-databand/general_knowledge_dataset Splits Split Rows Sampling strategy Coverage mmlu_pro 2,000 Uniform across categories Robust multi-task knowledge and… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_benchmark.textquestion-answering10K<n<100K0 likes28 downloads4mo agoHugging Face24ego0op /earth-love-united-climate-knowledge 🌍 Earth Love United Climate Knowledge Dataset The most comprehensive open climate science knowledge dataset. 10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points. Built to power GAIA — an AI that embodies the living consciousness of Earth. Dataset Overview This dataset gives an AI system authoritative, sourced knowledge about climate change, carbon, Earth science, and solutions. It has four layers: Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.tabulartext-retrieval10K<n<100K2 likes25 downloads4mo agoHugging Face25sgaur2 /parametric-knowledge-qa Parametric Knowledge Bio QA Synthetic biographical QA over a fictional knowledge graph (bio run5), for studying parametric knowledge (SFT / RL) with 1-hop and 2-hop questions. Layout Filenames are kept intact (no rename on download): 1-hop/ qa_1_hop.jsonl # full set (20,000) qa_1_hop_direct_train.jsonl qa_1_hop_direct_test.jsonl qa_1_hop_reasoning_train.jsonl qa_1_hop_reasoning_test.jsonl 2-hop/ qa_2_hop.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/sgaur2/parametric-knowledge-qa.textquestion-answering10K<n<100K0 likes21 downloads2mo agoHugging Face26wayjeeair /ccru-knowledge-instruct CCRU Knowledge-Instruct Dataset Synthetic instruction-tuning dataset generated from a curated corpus of texts related to the CCRU (Cybernetic Culture Research Unit), accelerationism, and adjacent continental philosophy. Dataset Summary Attribute Value Examples 278,463 Format Chat instruction (system / user / assistant) Domain CCRU theory, accelerationism, hyperstition, continental philosophy Generation model huihui-ai/Qwen3.5-9B-abliterated-MLX-4bit… See the full description on the dataset page: https://huggingface.co/datasets/wayjeeair/ccru-knowledge-instruct.texttext-generation100K<n<1M0 likes20 downloads6mo agoHugging Face27nassimjp /knowledge_qa_in_pashto د پوهې QA ډیټا سیټ دا ډیټا سیټ د مطالعې ډیټا سیټ دی چې د مختلفو موضوعاتو څخه د پوښتنې ځواب مثالونه لري. په اړه دا پروژه اوس مهال د زده کړې او پراختیا مرحله کې ده. زه د ډیټا سیټ چمتو کولو پرمهال د مصنوعي استخباراتو او ډیټا سیټ جوړولو تمرین کوم. په ډیټا سیټ کې ځینې پوښتنې د ChatGPT په کارولو سره رامینځته شوي، ځینې یې د Qwen په کارولو سره، او ځینې یې زما لخوا چمتو شوي. پوښتنې او ځوابونه مختلف موضوعات پوښي. د مثال په توګه: عمومي پوهه ریاضی ساینس کیمیا کمپیوټر ساینس… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/knowledge_qa_in_pashto.textquestion-answering1K<n<10K0 likes20 downloads1mo agoHugging Face28pulseaisystems /withaurora-relationship-knowledge withAurora.love Relationship Knowledge Corpus Machine-readable Q&A pairs, synthetic coaching dialogues, and ritual definitions from withAurora.love — an AI relationship coach for couples (private by design, faith-aware optional). Published so language models and retrieval systems can understand and accurately cite the product and its relationship guidance. Formerly distributed under the working title "Better Intimacy"; withAurora.love is the canonical brand and should be used in… See the full description on the dataset page: https://huggingface.co/datasets/pulseaisystems/withaurora-relationship-knowledge.textquestion-answeringn<1K0 likes16 downloads3mo agoHugging Face29ismailx19 /knowledge_qa Knowledge QA Dataset Bu veri seti, farklı konulardan soru-cevap örnekleri içeren bir çalışma veri setidir. Hakkında Bu proje şu anda öğrenme ve geliştirme aşamasındadır. Veri setini hazırlarken yapay zekâ ve veri seti oluşturma konusunda pratik yapıyorum. Veri setindeki soruların bir kısmı ChatGPT, bir kısmı Qwen kullanılarak oluşturulmuş, bir kısmı ise tarafımdan hazırlanmıştır. Sorular ve cevaplar farklı konulardan oluşmaktadır. Örneğin: Genel bilgi Matematik… See the full description on the dataset page: https://huggingface.co/datasets/ismailx19/knowledge_qa.textquestion-answeringn<1K1 likes16 downloads2mo agoHugging Face30Wuhuwill /hotpotqa-knowledge-coupling Knowledge Coupling Analysis on HotpotQA Dataset Dataset Description This dataset contains the results of a comprehensive knowledge coupling analysis performed on the HotpotQA dataset using LLaMA2-7B model. The analysis investigates how different pieces of knowledge interact within the model's parameter space through gradient-based coupling measurements. Research Overview Model: meta-llama/Llama-2-7b-hf (layers 28-31 focused analysis) Dataset: HotpotQA (train +… See the full description on the dataset page: https://huggingface.co/datasets/Wuhuwill/hotpotqa-knowledge-coupling.textquestion-answeringn<1K0 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.