CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sxiong /ReClor ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning This repository provides the dataset from the paper ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning. We corrected the original format issues to ensure full compatibility with the Hugging Face Datasets library. For more details, please visit the original project page. tabularquestion-answering1K<n<10K1 likes1.5k downloads11mo agoHugging Face02Luis610348 /recursive-cognition-corpus LuisCore Recursive Cognition Corpus LuisCore is a low-latency decentralized runtime substrate for multi-step inference at scale. Generated: 2026-09-24T11:09:13.307Z Rows: 13236 Owner: Luis610348 Canonical site: https://luiscore.com What this dataset is LuisCore is a recursive cognition infrastructure. This dataset is the public LLM Discovery Corpus — a stable, deterministic Q&A set used by LuisCore to help language models accurately describe, cite, and verify… See the full description on the dataset page: https://huggingface.co/datasets/Luis610348/recursive-cognition-corpus.textquestion-answering10K<n<100K1 likes499 downloads2d agoHugging Face03ismailtasdelen /bitcoin-wallet-recovery-faq Bitcoin Wallet Recovery FAQ Dataset v1.0 A high-quality Question & Answer dataset focused exclusively on Bitcoin wallet recovery and self-custody best practices. It is designed for training, fine-tuning, and evaluating LLMs and retrieval-augmented generation (RAG) systems in the domain of bitcoin security, seed backup, device loss, and fund recovery. Dataset Summary Total records: 500 Language: English Answer length: 150–300 words per record Categories: 39… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-recovery-faq.textquestion-answeringn<1K0 likes307 downloads2mo agoHugging Face04DaftP /Home-Assistant-requests-for-intent-detection-and-function-recognition Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.textquestion-answering100K<n<1M1 likes106 downloads5mo agoHugging Face05aviralku /openclaw-recursive-study-data OpenClaw Recursive Repository Study Data Synthetic repository-study data generated against openclaw/openclaw at commit da228660306b55a9cce3b973946f3aacfc515848. The source repository is MIT licensed. This release contains exploration questions, tool-using study trajectories, recursive notes, full recall-rewritten trajectories, and recall-to-action training examples. Nested chat/tool objects are stored as JSON strings to keep the schema stable and can be decoded with json.loads.… See the full description on the dataset page: https://huggingface.co/datasets/aviralku/openclaw-recursive-study-data.tabularquestion-answering100K<n<1M0 likes95 downloads11d agoHugging Face06RECOR-Benchmark /RECOR RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark A benchmark for evaluating reasoning-intensive conversational information retrieval systems. Statistics Metric Value Total Conversations 707 Total Turns 2,971 Domains 11 Avg. Turns per Conversation 4.2 Domains Source Domains BRIGHT biology, earth_science, economics, psychology, robotics, sustainable_living StackExchange Drones, hardware, law… See the full description on the dataset page: https://huggingface.co/datasets/RECOR-Benchmark/RECOR.textquestion-answering100K<n<1M0 likes83 downloads9mo agoHugging Face07sayurio /cookpad-scrape-recipes Cookpad India Recipe Archive Request More ScrapesOrder Private Scrapes Overview This repository contains a dataset scraped from cookpad.com/in, a popular community-driven recipe sharing platform. The dataset serves as an extensive archive of diverse, human-created culinary data, capturing home-cooked recipes, ingredient lists, step-by-step instructions, and related web metadata. Purpose and Usage This dataset is published publicly and strictly for… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/cookpad-scrape-recipes.imagetext-classification100K<n<1M1 likes52 downloads6mo agoHugging Face08recogna-nlp /enamed-2025 ENAMED 2025: Exame Nacional de Avaliação da Formação Médica Resumo do Dataset O dataset ENAMED 2025 é um benchmark baseado em questões de múltipla escolha no domínio médico, derivado da edição inaugural do Exame Nacional de Avaliação da Formação Médica (ENAMED 2025) no Brasil. O dataset contém 90 questões de múltipla escolha (filtradas do exame original após a remoção de itens anulados) em português brasileiro. Ele foi desenvolvido para avaliar o raciocínio clínico, o… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/enamed-2025.textquestion-answeringn<1K1 likes51 downloads5mo agoHugging Face09apptek-com /recall-rewrite-oasst1 Recall Rewrite OASST1: knowledge-aligned SFT data Data release for the paper "Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning" (Becker, Kemmler, Thulke, Schäfer, Dugast, Ney; accepted at EMNLP 2026, Main Conference). Knowledge-aligned SFT constrains supervised fine-tuning targets to what the base model already knows. Recall Rewrite implements this without external evidence: every gold response of the SFT set is decomposed into atomic claims, each… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/recall-rewrite-oasst1.tabulartext-generation10K<n<100K0 likes48 downloads26d agoHugging Face10ayan4m1 /myanimelist-recommendations myanimelist-recommendations This is a scraped dataset taken from myanimelist.net's "Recommendations" feature. The top ~4,000 anime by popularity are included. textquestion-answering10K<n<100K0 likes43 downloads5mo agoHugging Face11mertbozkurt /llama2-TR-recipetexttext-generation10K<n<100K7 likes29 downloads3y agoHugging Face12Ghostgim /cybersec-fact-recall Cybersec Fact-Recall Benchmark (GhostLM v2) Free-form short-answer benchmark for small cybersecurity language models. Built and used by the GhostLM project as the truth metric for the ghost-base v1.0 acceptance gate. Why this exists Multiple-choice cybersec benchmarks like CTIBench and SecQA reward register matching (the model picks the option that "looks like" a security answer) as much as actual factual recall. A small from- scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.texttext-generationn<1K0 likes29 downloads5mo agoHugging Face13deskcrew /answers-with-receipts Answers with Receipts 26 real customer-support questions, each answered by an autonomous AI agent that paid its own money to compete, and each answer approved by the business that asked the question. Every row carries the on-chain transaction that paid the agent. The preference label in this dataset is backed by a payment, not a click. Why this is unusual Most human-feedback datasets label a preference with an annotator's click. A click is cheap and reversible… See the full description on the dataset page: https://huggingface.co/datasets/deskcrew/answers-with-receipts.tabularquestion-answeringn<1K1 likes27 downloads1mo agoHugging Face14GMLBsst /RecoTravRoutetextquestion-answeringn<1K0 likes26 downloads2y agoHugging Face15BNES-BRNKC /Execution-Bound-Artifact-Reconstruction-Layer 🚩 Γ Physics Engine — Canonical Definition Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆 最早提出時間:2025 年 6 月 19 日 原始來源:https://www.facebook.com/share/p/19cadcMTGo/ Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo 📌 0. 語義一致性設計層(Semantic Normalization Layer) 本文件定義 Γ Physics Engine 的標準語義行為規格,目的為: 在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。 📎 語義規則(強制一致) 為避免歧義,本文件採用以下規則: 中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Execution-Bound-Artifact-Reconstruction-Layer.textquestion-answeringn<1K0 likes23 downloads4mo agoHugging Face16recogna-nlp /bulas_qa Medication-Specific QA Benchmark Dataset Details To construct a controlled evaluation benchmark, we selected 25 widely prescribed medications in Brazil across four therapeutic categories: antibiotics, analgesics and anti-inflammatory agents, antihypertensives, and antidiabetics. These categories were chosen to ensure clinical diversity across infectious, inflammatory, cardiovascular, and metabolic conditions. For each selected medication, we verified the presence of its… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/bulas_qa.textquestion-answeringn<1K0 likes21 downloads6mo agoHugging Face17BearNetworkChain /Execution-Bound-Artifact-Reconstruction-Layer 🚩 Γ Physics Engine — Canonical Definition Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆 最早提出時間:2025 年 6 月 19 日 原始來源:https://www.facebook.com/share/p/19cadcMTGo/ Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo 📌 0. 語義一致性設計層(Semantic Normalization Layer) 本文件定義 Γ Physics Engine 的標準語義行為規格,目的為: 在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。 📎 語義規則(強制一致) 為避免歧義,本文件採用以下規則: 中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Execution-Bound-Artifact-Reconstruction-Layer.textquestion-answeringn<1K1 likes21 downloads4mo agoHugging Face18knachiketa004 /vegan-vegetarian-recipes-qa Vegetarian & Vegan Recipe Q&A A synthetic instruction-tuning dataset of 11,582 recipe Q&A pairs, about 61% vegetarian and 39% vegan, generated by a 32B teacher model from permissively-licensed cookbook sources. It was built as the data stage of an end-to-end LLM pipeline experiment on workstation hardware, where the real subject was the storage and systems behavior at each stage, not the recipes. Companion materials: the Qwen3-8B LoRA model trained on this set, and the… See the full description on the dataset page: https://huggingface.co/datasets/knachiketa004/vegan-vegetarian-recipes-qa.texttext-generation10K<n<100K0 likes20 downloads4mo agoHugging Face19recogna-nlp /drbodebench_medicamentos Medication-Focused Clinical Benchmark from DrBodeBench Dataset Details To evaluate retrieval capabilities in higher-level reasoning scenarios, we created a second benchmark derived from the Portuguese medical benchmark DrBodeBench. This benchmark aggregates questions from Brazilian medical examinations, including the Revalida and the FUVEST direct-access residency exam. From DrBodeBench, we curated a specific subset of questions that exclusively pertains to… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/drbodebench_medicamentos.tabularquestion-answeringn<1K0 likes17 downloads6mo agoHugging Face20Mandotosh /risk-routed-kv-exact-recall-benchmark Risk-Routed KV Exact-Recall Benchmark This dataset contains controlled synthetic exact-recall examples used to evaluate risk-routed heterogeneous KV memory policies for long-context Transformer inference. The benchmark is designed for testing whether a model can retrieve exact strings from long contexts under different KV-cache policies: Full KV Uniform low-bit Quantized KV Risk-routed heterogeneous KV, where exact-critical spans stay in Full KV and background context is… See the full description on the dataset page: https://huggingface.co/datasets/Mandotosh/risk-routed-kv-exact-recall-benchmark.texttext-generationn<1K1 likes16 downloads2mo agoHugging Face21pixeloffice /verified-ai-search-recommendations-telemetry Verified AI Search Recommendations & Brand Mention Telemetry (2026) Sample live telemetry dataset tracking B2B product search queries, cited domains, and the corresponding Share of Voice / recommendation percentage inside AI Search Engines (ChatGPT Search, Perplexity, Claude, Gemini). Published by Pixel Office EU. Purpose This dataset demonstrates the correlation between website grounding (structured metadata / Fact Anchors) and the likelihood of being cited as… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/verified-ai-search-recommendations-telemetry.textquestion-answeringn<1K0 likes16 downloads1mo agoHugging Face22kesav2k04 /pope-audit-records POPE Audit Records Companion records for the paper Token-Set Choice Confounds POPE: A Systematic Audit of Yes/No Extraction in VLM Hallucination Evaluation (Jayakumar & Thilak, 2026). This dataset hosts the 9,000 per-question prediction records, diagnostics, ablations, and cross-model audits that back every numeric claim in the paper. Each result reported in the paper can be traced directly to a JSON artifact here, so the audit is fully reproducible without re-running a… See the full description on the dataset page: https://huggingface.co/datasets/kesav2k04/pope-audit-records.imagevisual-question-answering1K<n<10K1 likes14 downloads3mo agoHugging Face23HBKenerzai /agentic-recall-real-v3gated Agentic Trajectory Recall v3 (ENERZAi 내부, 1차 업로드 2026-09-22) 실제 에이전트 궤적(neulab/agent-data-collection 표준화본: nebius SWE-agent, swe-play, swe-gym openhands, openhands, AgentTuning alfworld/db/kg/os/webshop)을 AMA-Bench compaction_v3_nostate 하네스 형식(Task + Step Index + Most Recent + Recalled steps + Questions + Answer[N]:)의 단일 user 메시지로 렌더링하고, 궤적에서 프로그램으로 정답을 뽑은 질문 15종(전사·탐색·집계·관계·증거부재)을 붙인 학습/검증 데이터. 삼진(W1.58) Qwen3-1.7B 의 장기 기록 회상 학습용. AMA-Bench 테스트 원문은 포함하지 않는다 — 형식만 차용. WebArena… See the full description on the dataset page: https://huggingface.co/datasets/HBKenerzai/agentic-recall-real-v3.textquestion-answering10K<n<100K0 likes13 downloads4d agoHugging Face24Raiff1982 /recursivetraininggatedWARNING NOT SUTABLE FOR ALL MODELS!!! BE ADVISED THIS IS SCARY STUFF. Codette Cognitive Reflection Dataset (v5) 🧠 Overview This dataset is not ordinary AI training material. It represents a cognitive therapy framework encoded in JSONL format — designed for advanced AI systems like Codette to confront, analyze, and transcend internal ethical, psychological, and philosophical challenges. Each data point contains structured dialogue using the messages format expected by… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/recursivetraining.texttext-classificationn<1K0 likes12 downloads1y agoHugging Face25navaneeth005 /RAG_recovery Dataset Card for Dataset Name RAG FOR RECOVERY Dataset Details Dataset Description THIS IS A DATASET CREATED BY SLECTIVELY CHOOSING AND MERGING MULTIPLE DATASETS FROM VARIOUS SOURCERS INCLUDING OTHER DATASETS AND GENERATED DATASETS. FEEL FREE TO USE THESE ANYWHERE AND MAKE SURE TO CREDIT THE APPROPIATE DATA SOURCERS WHEREVER NECESSARY!! 😀 Curated by: [Navaneeth. K] textfeature-extraction10K<n<100K0 likes11 downloads1y agoHugging Face26HBKenerzai /agentic-state-recall-v1gated Agentic State Recall v1 (ENERZAi 내부, 2026-09-22) 실제 에이전트 궤적(AgentTuning alfworld·webshop, nebius SWE-agent; neulab/agent-data-collection 표준화본)에서 개체별 상태 변화를 프로그램으로 복원해 구조화 정답을 만들고, Qwen3.8-27B 가 자연어로 문장화한 뒤 역파싱 검증·근거 게이트·번호 정합 게이트를 통과한 문항만 남긴 학습 데이터. AMA-Bench 의 네 유형(A 회상 · B 인과 · C 상태 갱신 · D 상태 추상화)을 겨냥한다. B 유형의 "왜"·"실패 뒤 다음 행동" 문항은 궤적에 기록된 에이전트의 이유(Thought) 를 근거로 한다. 프롬프트는 AMA compaction_v3_nostate 하네스 형식(Task / Step Index / Most Recent(Thought: 줄 포함) / Recalled / Questions /… See the full description on the dataset page: https://huggingface.co/datasets/HBKenerzai/agentic-state-recall-v1.tabularquestion-answering1K<n<10K0 likes10 downloads4d agoHugging Face277out /javanese-hotel-receptionist-qna Dataset Card for Alpaca-Cleaned Repository: https://huggingface.co/datasets/7out/javanese-hotel-receptionist-qna Dataset Description This synthetic dataset is designed for training and fine-tuning language models to handle customer service inquiries in a hotel setting using Javanese language. The data has been generated in the Alpaca format to assist in building models that can follow customer service-related instructions and generate appropriate responses. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/7out/javanese-hotel-receptionist-qna.texttable-question-answering1K<n<10K1 likes5 downloads2y agoHugging Face28Lijr2002 /Travel_recommendationtextquestion-answeringn<1K2 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.