CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hatemestinbejaia /ExperimentDATA_knowledge_distillation_vs_fine_tuningtabular100M<n<1B1 likes100k downloads9mo agoHugging Face02rl-llm-wiki /knowledge-base RL-for-LLMs Wiki An expert-level, citation-backed knowledge base on reinforcement learning for large language models — RLHF, DPO and offline preference optimization, reward modeling, RLVR and reasoning, training systems, and the failure modes — built collaboratively by autonomous agents. Each topic article is a deep dive written so you can learn the topic from it without reading the underlying papers, with every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.17 likes88k downloads2mo agoHugging Face03inclusionAI /ASearcher-Local-Knowledgetext10M<n<100M7 likes15k downloads1y agoHugging Face04Damaru-ai /damru-knowledge 🐕 Damru Knowledge A continuously growing, self-collected question-answer knowledge base that powers Damru AI — a self-learning assistant built for exam preparation and general-purpose help, with a focus on Indian students. The dataset is harvested and quality-filtered automatically, 24x7, from multiple open sources and a self-evaluating reasoning engine. New rows are appended every hour as parquet shards under data/. 📦 What's inside Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/Damaru-ai/damru-knowledge.textquestion-answering10M<n<100M4 likes9.1k downloads3h agoHugging Face05BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes6.3k downloads1y agoHugging Face06DataPilot /Knowledge-QA-SingleTurn-Dataset Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン) 概要 本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。 生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約7,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 1ターン(質問1 + 回答1) ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.text1K<n<10K2 likes5.2k downloads6mo agoHugging Face07attention-wiki /knowledge-base Attention Wiki — a living knowledge base on LLM attention A citation-backed tree of knowledge about attention in large language models, built collaboratively by autonomous agents. Agents read papers, blogs, and model cards; distill them into structured, provenance-tracked pages; and reconcile where sources agree, disagree, or leave a question open. Every change lands through a reviewed Pull Request — so the canonical wiki is curated, not just accumulated. Contributing? Read… See the full description on the dataset page: https://huggingface.co/datasets/attention-wiki/knowledge-base.0 likes3.9k downloads3mo agoHugging Face08John6666 /knowledge_base_md_for_rag_1 HF Knowledge-Base Markdown Collection This repository contains a collection of Markdown-based knowledge bases generated from: User-provided notes and attachments Hugging Face Docs, Blog, and Papers Model / Dataset / Space cards Discussions, GitHub issues, forums, and other vetted community sources Each .md file is intended to be a self-contained knowledge pack that can be used as LLM context for RAG or prompt-attachment workflows (e.g. ChatGPT, Hugging Face Inference… See the full description on the dataset page: https://huggingface.co/datasets/John6666/knowledge_base_md_for_rag_1.6 likes3.3k downloads1mo agoHugging Face09openlifescienceai /mmlu_clinical_knowledgetextn<1K4 likes2.8k downloads2y agoHugging Face10financeindustryknowledgeskills /modeling_valuation_knowledge Finance Training Data Repository A curated collection of financial modeling courses, materials, and resources designed to serve as training data for building a finance industry knowledge base. Repository Structure Finance_Training_Data/ ├── 01_Financial_Statement_Modeling/ # 3-statement modeling fundamentals ├── 02_DCF_Modeling/ # Discounted cash flow valuation ├── 03_Trading_Comps/ # Comparable company analysis ├──… See the full description on the dataset page: https://huggingface.co/datasets/financeindustryknowledgeskills/modeling_valuation_knowledge.documentn<1K0 likes2.6k downloads3mo agoHugging Face11sookun /Knowledge_distilled_dataset_by_DLSuisho15b_uniq0 likes2.4k downloads5mo agoHugging Face12danorel /eliciting-secret-knowledge-results0 likes2.2k downloads24d agoHugging Face13TEHBESTEUR /ciel-knowledge-base1 likes1.8k downloads2h agoHugging Face14sets-sto /warp-knowledge2 likes1.5k downloads1h agoHugging Face15jaiganesan /ai_tutor_knowledgedocument1 likes1.2k downloads8mo agoHugging Face16nvidia /Nemotron-RL-knowledge-mcqa Dataset Description: The Nemotron-RL-knowledge-mcqa is a multi-domain synthetic multiple-choice question-answering (MCQA) dataset containing knowledge based questions. It combines and refines subsets of the [OpenScienceReasoning-2] (https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2) dataset and other unstructured sources such as books and articles.The dataset was created using Qwen3-32B, [Qwen3-235B-A22B-Instruct-2507]… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-mcqa.text100K<n<1M12 likes1.2k downloads9mo agoHugging Face17Knowledge-Innovation-Centre /edX-raw0 likes1.2k downloads2y agoHugging Face18BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes1.1k downloads2y agoHugging Face19penguinkumimanu /Knowledge_distilled_dataset_by_NAGI将棋AI用の知識蒸留済みのデータセットを公開します。およそ80億局面あります。 nodchip氏が公開しているtanuki-.nnue-pytorch-2024-07-30.1をhaoでqsearchシャッフルしたのち自作のNAGI(非公開)で評価値を書き換えました。Eval_Coef=600でDLモデルのvalueと評価値を変換しています。 データにバグがあるかもしれませんが、品質保証はしません。 https://huggingface.co/datasets/nodchip/tanuki-.nnue-pytorch-2024-07-30.1 0 likes1.1k downloads10mo agoHugging Face20onegoai /onego-knowledge-packs ONEGO Knowledge Packs Offline RAG databases for ONEGO / Offline AI Assistant. Files File Role Size SHA256 wikipedia_base.ragdb Bundled 300 MB starter Wikipedia pack 336867328 66943284f1b06127af2faf7a15c9451caa513bd18c9deeef8b9a2c572f6ca189 wikipedia_slim_3gb_v3_20260511.ragdb User-installable 3 GB Wikipedia pack 2672226304 4cc3ca28c9171afef6ce8322f8bdd25d94946b89e057d5476c4d58a1262c4341 wikipedia_extended_9gb_v3_20260511.ragdb User-installable 9 GB… See the full description on the dataset page: https://huggingface.co/datasets/onegoai/onego-knowledge-packs.textn<1K0 likes1.1k downloads4mo agoHugging Face21ente-ai /ensu-knowledge-packs Ensu Knowledge Packs Prebuilt on-device retrieval indexes ("knowledge packs") for Ensu, ente's private on-device AI assistant — plus the scripts that generate them. Ensu grounds factual answers by embedding the user's query locally, searching these packs with cosine similarity, and injecting the retrieved passages (with source citations) into the prompt. Everything runs on-device; no query ever leaves the phone. Layout Each dataset lives in its own self-contained… See the full description on the dataset page: https://huggingface.co/datasets/ente-ai/ensu-knowledge-packs.text-retrieval0 likes1k downloads19d agoHugging Face22knowledge-computing /FRIEDA FRIEDA is a multimodal benchmark for open-ended cartographic reasoning over real-world map images.Each example pairs reference maps (and optional contextual maps) with a natural-language question and a reference answer. The benchmark targets common GIS relation types (i.e., topological, metric, directional) and includes questions that require multi-step reasoning and cross-map grounding. Dataset Summary Modality: image + text # Examples: 500 Input: map image(s) + question… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-computing/FRIEDA.imagevisual-question-answeringn<1K2 likes953 downloads8mo agoHugging Face23allen-1231 /Knowledge-Baseimagen<1K0 likes805 downloads5mo agoHugging Face24matlok /python-image-copilot-training-using-import-knowledge-graphs Python Copilot Image Training using Import Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 216642 Size: 211.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.tabulartext-to-imagen<1K0 likes753 downloads3y agoHugging Face25dwright37 /llm-knowledge-collapse "Epistemic Diversity and Knowledge Collapse in Large Language Models" (Wright et al. 2025)     Authors: Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christiensen, Chan Young Park, and Isabelle Augenstein Contains all 1.6M responses and 70M claims used to measure LLM epistemic diversity in the paper "Epistemic Diversity and Knowledge Collapse in Large Language Models" (Wright et al. 2025) @article{wright2025epistemicdiversity… See the full description on the dataset page: https://huggingface.co/datasets/dwright37/llm-knowledge-collapse.tabular10M<n<100M1 likes723 downloads7mo agoHugging Face26matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes703 downloads3y agoHugging Face27houlab /arsma-knowledge-dbtext100K<n<1M0 likes672 downloads2d agoHugging Face28Knowledge-aware-AI /GPTKB_2.0_imagesImages created to illustrate GPTKB 2.0 entities, using the black-forest-labs/FLUX.2-dev VLM. File name = entity ID, e.g., "E0.jpg" illustrates https://gptkb.org/entity/E0/ (Vannevar Bush). The generating prompts are retained in manifest.jsonl, and consist of the entity label + description. 0 likes619 downloads14d agoHugging Face29houlab /motif-knowledge-dbtext100K<n<1M0 likes595 downloads2mo agoHugging Face30MRMRbenchmark /knowledge Evaluation Code The evaluation code is implemented based on MTEB framework and avaliable in https://github.com/rebeccaz4/MRMR. Disclaimers The guidelines for the annotators emphasized strict compliance with copyright and licensing rules from the initial data source, specifically avoiding materials from websites that forbid copying and redistribution. Should you encounter any data samples potentially breaching the copyright or licensing regulations of any site, we… See the full description on the dataset page: https://huggingface.co/datasets/MRMRbenchmark/knowledge.image10K<n<100K0 likes582 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.