CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Damaru-ai /damru-knowledge 🐕 Damru Knowledge A continuously growing, self-collected question-answer knowledge base that powers Damru AI — a self-learning assistant built for exam preparation and general-purpose help, with a focus on Indian students. The dataset is harvested and quality-filtered automatically, 24x7, from multiple open sources and a self-evaluating reasoning engine. New rows are appended every hour as parquet shards under data/. 📦 What's inside Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/Damaru-ai/damru-knowledge.textquestion-answering10M<n<100M4 likes9.5k downloads26m agoHugging Face02BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes6.4k downloads1y agoHugging Face03openlifescienceai /mmlu_clinical_knowledgetextn<1K4 likes2.6k downloads2y agoHugging Face04nvidia /Nemotron-RL-knowledge-mcqa Dataset Description: The Nemotron-RL-knowledge-mcqa is a multi-domain synthetic multiple-choice question-answering (MCQA) dataset containing knowledge based questions. It combines and refines subsets of the [OpenScienceReasoning-2] (https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2) dataset and other unstructured sources such as books and articles.The dataset was created using Qwen3-32B, [Qwen3-235B-A22B-Instruct-2507]… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-mcqa.text100K<n<1M12 likes1.5k downloads10mo agoHugging Face05BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes1.1k downloads2y agoHugging Face06knowledge-computing /FRIEDA FRIEDA is a multimodal benchmark for open-ended cartographic reasoning over real-world map images.Each example pairs reference maps (and optional contextual maps) with a natural-language question and a reference answer. The benchmark targets common GIS relation types (i.e., topological, metric, directional) and includes questions that require multi-step reasoning and cross-map grounding. Dataset Summary Modality: image + text # Examples: 500 Input: map image(s) + question… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-computing/FRIEDA.imagevisual-question-answeringn<1K2 likes992 downloads8mo agoHugging Face07matlok /python-image-copilot-training-using-import-knowledge-graphs Python Copilot Image Training using Import Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 216642 Size: 211.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.tabulartext-to-imagen<1K0 likes773 downloads3y agoHugging Face08matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes713 downloads3y agoHugging Face09dwright37 /llm-knowledge-collapse "Epistemic Diversity and Knowledge Collapse in Large Language Models" (Wright et al. 2025)     Authors: Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christiensen, Chan Young Park, and Isabelle Augenstein Contains all 1.6M responses and 70M claims used to measure LLM epistemic diversity in the paper "Epistemic Diversity and Knowledge Collapse in Large Language Models" (Wright et al. 2025) @article{wright2025epistemicdiversity… See the full description on the dataset page: https://huggingface.co/datasets/dwright37/llm-knowledge-collapse.tabular10M<n<100M1 likes666 downloads7mo agoHugging Face10houlab /arsma-knowledge-dbtext100K<n<1M0 likes654 downloads3d agoHugging Face11MRMRbenchmark /knowledge Evaluation Code The evaluation code is implemented based on MTEB framework and avaliable in https://github.com/rebeccaz4/MRMR. Disclaimers The guidelines for the annotators emphasized strict compliance with copyright and licensing rules from the initial data source, specifically avoiding materials from websites that forbid copying and redistribution. Should you encounter any data samples potentially breaching the copyright or licensing regulations of any site, we… See the full description on the dataset page: https://huggingface.co/datasets/MRMRbenchmark/knowledge.image10K<n<100K0 likes584 downloads10mo agoHugging Face12houlab /motif-knowledge-dbtext100K<n<1M0 likes576 downloads2mo agoHugging Face13hoanganhpham /HealthBench_Knowledge_Question_250Ktext100K<n<1M0 likes499 downloads1y agoHugging Face14matlok /python-audio-copilot-training-using-function-knowledge-graphs Python Copilot Audio Training using Global Functions with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.tabulartext-to-audion<1K1 likes471 downloads3y agoHugging Face15matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes461 downloads3y agoHugging Face16matlok /python-audio-copilot-training-using-import-knowledge-graphs Python Copilot Audio Training using Imports with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.tabulartext-to-audion<1K0 likes453 downloads3y agoHugging Face17knowledge-in-visual-synthesis /v1 Knowledge in Visual Synthesis This dataset contains prompt–image examples for evaluating and studying knowledge-intensive visual synthesis. Samples are organized by contributor as dataset subsets (configs), with each upload version exposed as a split. Dataset structure Subset Splits byx v1, v2 yuner v1, v2 zanyi v1, v2, v3 jiayu v1, v2, v3 sherry v1, v2 yujunz v1 The byx/v1 split contains 140 unique prompts and 300 generated images. For… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.image1K<n<10K0 likes396 downloads1mo agoHugging Face18matlok /python-audio-copilot-training-using-inheritance-knowledge-graphs Python Copilot Audio Training using Inheritance and Polymorphism Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each base class for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-inheritance-knowledge-graphs.tabulartext-to-audion<1K0 likes379 downloads3y agoHugging Face19MMB-25 /knowledgeimage10K<n<100K0 likes378 downloads1y agoHugging Face20matlok /python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-audion<1K0 likes371 downloads3y agoHugging Face21matlok /python-image-copilot-training-using-inheritance-knowledge-graphs Python Copilot Image Training using Inheritance and Polymorphism Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 259017 Size: 135.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: {… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-inheritance-knowledge-graphs.tabulartext-to-imagen<1K0 likes337 downloads3y agoHugging Face22PreFLMR /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KRtext10M<n<100M0 likes313 downloads2y agoHugging Face23matlok /python-image-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312836 Size: 294.1 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-imagen<1K0 likes280 downloads3y agoHugging Face24matlok /python-image-copilot-training-using-function-knowledge-graphs Python Copilot Image Training using Function Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 134357 Size: 130.5 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-function-knowledge-graphs.tabulartext-to-imagen<1K0 likes275 downloads3y agoHugging Face25ranjithraj /cancer-knowledge-base Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation The only open CC-BY-4.0 oncology knowledge base that combines: 110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this. A provable 152-question MCQ benchmark — every answer derives from this KB's own structured data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.tabularquestion-answering10K<n<100K0 likes260 downloads1mo agoHugging Face26qleap /Knowledge_distilled_dataset_by_NAGISA_V3 NAGISA_V3 teacher shards Training data distilled from self-play of the search engine attic reading the NNUE weights NAGISA_V3, in the shape the trainers read directly. Starting from a balanced-opening book, the engine played itself at MultiPV=5, recording the root score and the candidate moves at every ply; manaka-teacher turned that corpus into raw MPK1 streams, and manaka-pack folded identical positions into one row each and wrote these parquet shards. Rows: 280,621,202… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Knowledge_distilled_dataset_by_NAGISA_V3.tabularreinforcement-learning100M<n<1B0 likes203 downloads21d agoHugging Face27Knowledge-aware-AI /GPTKB_v2 GPTKB 2.0 Browse, query or ask the full knowledge base at gptkb.org. Two artifact lines, one knowledge base (snapshot gpt51_270226, 38,450,135 source facts): File What it is gptkb_v2.0.1_rdf.nt.gz Simplified RDF export. Use this one. gptkb_v2.0.1_full_record/ The full six-term record, as Parquet. gptkb_v2.0_rdf.nt.gz Superseded interim RDF export, kept for citability. Simplified RDF export (gptkb_v2.0.1_rdf.nt.gz) Format N-Triples… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v2.text10M<n<100M0 likes192 downloads21d agoHugging Face28nvidia /Nemotron-RL-knowledge-openqa Dataset Description: The Nemotron-RL-knowledge-openQA is a multi-domain synthetic dataset containing knowledge based questions. It is built from unstructured sources such as books and articles and consists of question–answer pairs requiring short responses. The dataset covers a wide range of domains, including physics, biology, mathematics, computer science, engineering, chemistry, law, and others. This dataset is released as part of NVIDIA NeMo Gym, a framework for building… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-openqa.text100K<n<1M10 likes184 downloads10mo agoHugging Face29jiosephlee /auxiliary-views-knowledge-acquisition Auxiliary Views Knowledge Acquisition This repository contains the cleaned source documents and evaluation probes used in Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views (arXiv:2609.04180). News August 21, 2026: Our paper was accepted to Findings of EMNLP 2026. Configurations Configuration Split Rows documents train 30 factual_cloze test 6,435 factual_mcqa_5shot test 4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.texttext-generation10K<n<100K1 likes182 downloads13d agoHugging Face30open-athena /nemotron-gym-knowledge-openqa-v2-qwen3.5-122b-32k-tracestext1K<n<10K0 likes159 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.