CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01domenicrosati /TruthfulQA Dataset Card for TruthfulQA Dataset Summary TruthfulQA: Measuring How Models Mimic Human Falsehoods We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.textquestion-answeringn<1K52 likes4.5k downloads4y agoHugging Face02SamuelChien821 /blobfish-domainbench-24 Blobfish DomainBench-24 v3.3.2 DomainBench-24 v3.3.2 is the public release record for 24 realistic stateful agent tasks across six professional domains. Every task runs in a checked-in SQLite-backed MCP world with typed read/write tools and deterministic state, exact argument-aware trace, containment, persisted causal workpaper, stakeholder handoff, provider-native exact-record readback for every changed domain row, and persisted-result readback verification. The release… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24.documentquestion-answeringn<1K1 likes583 downloads24d agoHugging Face03HYdsl /Open-domain_Financial_QA Large-scale Open-domain Financial QA (LOFin) This repository accompanies the dataset used in the paper: Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents (ACL 2025 Findings) We introduce a benchmark for open-domain financial question answering over standardized documents, with a focus on multi-document and multi-hop reasoning. 📊 Benchmark Overview This benchmark targets realistic financial QA tasks… See the full description on the dataset page: https://huggingface.co/datasets/HYdsl/Open-domain_Financial_QA.question-answering1K<n<10K0 likes479 downloads4mo agoHugging Face04domofon /Domofon-Cot-Conversations-700k Domofon-Cot-Conversations-700k Synthetic XML conversation data for training small language models on reasoning, instruction following, XML formatting, and tool-use traces. Repository: domofon/Domofon-Cot-Conversations-700k What is inside The dataset contains cleaned generated XML conversations from six families: conv: multi-turn factual conversations with tool-use traces. instruct: text-processing instructions, including deterministic count tool calls. ds:… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Domofon-Cot-Conversations-700k.tabulartext-generation1M<n<10M1 likes391 downloads4mo agoHugging Face05botcoinmoney /domain-agnostic-causal-reasoning-tuning Domain-Agnostic Causal Reasoning Tuning Dataset Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network. The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.textquestion-answering10K<n<100K1 likes164 downloads6mo agoHugging Face06dendriteholdings /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes132 downloads22d agoHugging Face07domofon /domofon-cot-part1 Domofon COT Dataset - Part 1 Mixed Chain-of-Thought датасет на русском языке. (Russian language) Статистика Строк | lines : 700,000 Токенов: | tokens 594,412,523 Модели-генераторы (Distill Generation Models used): Deepseek R1 671B Mistral Large Devstral Mini Qwen 480B Структура (Structure) user - вопрос пользователя (User Question) thinking - цепочка рассуждений (Chain-of-Thought) answer - финальный ответ (Final answer) Связанные… See the full description on the dataset page: https://huggingface.co/datasets/domofon/domofon-cot-part1.texttext-generation100K<n<1M2 likes122 downloads9mo agoHugging Face08prem-research /domains Domains dataset Documentation coming soon textquestion-answering10K<n<100K4 likes118 downloads2y agoHugging Face09bluecolor777 /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes82 downloads14d agoHugging Face10emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes75 downloads8d agoHugging Face11cellos /DomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset] This Dataset includes 980,065 geographic names as of September 10, 2023. It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories. Example: feature_name: Abercrombie Gulch GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.tabularquestion-answering100K<n<1M2 likes53 downloads3y agoHugging Face12MidGUI /Mid-Training_data_of_separate_domains Breaking the Data Barrier – Building GUI Agents Through Task Generalization This is the official dataset repository of GUIMid 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains Observation WebArena (PR) WebArena (SR) AndroidWorld (SR) GUI Post-Training Only Image 26.3 6.2 9.0 Public Baselines GPT-4o-2024-11-20 Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.texttext-generation1M<n<10M0 likes53 downloads1y agoHugging Face13somasekhar-dev /nexttoken-pmkisan-domain-sft-data NextToken pmkisan domain SFT data (v1) Grounded multilingual QA dataset for fine-tuning somasekhar-dev/NextToken-model-1 on the Indian government-schemes / banking-financial domain. Generated by a pipeline (chunk source docs -> generate questions -> generate grounded answers -> validate/assemble) using a local LLM generator, from ~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA, banking products, insurance, savings instruments, etc.). Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.tabularquestion-answering1K<n<10K0 likes53 downloads7d agoHugging Face14Yana /ft-llm-2026-domain-specific-qa FT-LLM 2026 Domain-Specific QA A Japanese financial-domain visual QA dataset used for Phase 3 domain fine-tuning of the COMPASS Vision-Language Model. Question–answer pairs were generated with Qwen3-VL from scraped Japanese government financial PDFs (Cabinet Office, Financial Services Agency, Ministry of Finance), covering four difficulty tiers: (A) numeric extraction, (B) rate-of-change & comparison, (C) financial formula application, and (D) complex reasoning. Each answer includes… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-domain-specific-qa.imagevisual-question-answering10K<n<100K0 likes51 downloads5mo agoHugging Face15lablup /tariff_trade_domain.synthetic_trade_qa_kr Korean Trade Domain QA Dataset A Korean-language question-answering dataset for the international trade domain, combining 21,399 QA pairs from three complementary sources: official 무역영어 1급 certification exam questions, trade terminology definitions, and lecture-derived QA pairs. Designed for fine-tuning and evaluating LLMs on Korean trade domain knowledge. Dataset Description The dataset covers vocabulary, regulations, procedures, and concepts from Korean… See the full description on the dataset page: https://huggingface.co/datasets/lablup/tariff_trade_domain.synthetic_trade_qa_kr.textquestion-answering10K<n<100K1 likes48 downloads4mo agoHugging Face16EnDevSols /Multi-Domain-Reasoning-SFT Multi-Domain-Reasoning-SFT Dataset Summary The Multi-Domain-Reasoning-SFT dataset by EnDevSols is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset designed to train large language models in deep reasoning, technical analysis, and complex problem-solving. Consisting of nearly 580,000 meticulously structured examples, this dataset is specifically engineered to teach models how to "think" before they answer. It separates the internal cognitive process… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/Multi-Domain-Reasoning-SFT.texttext-generation100K<n<1M0 likes45 downloads5mo agoHugging Face17Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes43 downloads16d agoHugging Face18Somtharu181coder /science_behavioral_and_domain_diversity_dataset Nepali Science SFT Dataset — Clean Candidate A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script. This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering. Dataset Overview Property Value Dataset file clean_candidate.jsonl Records 29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.texttext-generation10K<n<100K0 likes36 downloads1mo agoHugging Face19michsethowusu /Code-170k-dombe Dataset Description Code-170k-dombe is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Dombe, making coding education accessible to Dombe speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Dombe language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-dombe.texttext-generation100K<n<1M0 likes33 downloads11mo agoHugging Face20HienNGuit /Domain-ViWikiQA Vietnamese Domain WikiQA Dataset Dataset Overview Vietnamese WikiQA is a Vietnamese question-answering dataset built from Vietnamese Wikipedia. It is designed for QA research, dataset analysis, and difficulty modeling. The dataset contains 7,592 QA pairs after generation, automatic filtering, human verification, and release normalization. Dataset Structure Each row in the public release represents one QA instance generated from a Wikipedia section… See the full description on the dataset page: https://huggingface.co/datasets/HienNGuit/Domain-ViWikiQA.textquestion-answering1K<n<10K2 likes33 downloads4mo agoHugging Face21DomainLLM /gerlayqa-bgb-paraphrased GerLayQA-BGB Paraphrased 🇩🇪⚖️ Dataset Description This is a paraphrased and restructured version of the GerLayQA BGB (Bürgerliches Gesetzbuch / German Civil Code) dataset, specifically prepared for fine-tuning large language models on German civil law question-answering tasks. Key Features 5,255 high-quality QA pairs about German Civil Law (BGB) Paraphrased questions to remove plagiarism while maintaining legal accuracy Structured 7-section answers following… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-bgb-paraphrased.textquestion-answering1K<n<10K0 likes32 downloads1y agoHugging Face22domofon /domofon-cot-part2 Domofon COT Dataset - Part 2 (Mixed Styles) Mixed Chain-of-Thought датасет на русском языке с разными стилями генерации. Статистика Строк: 321,072 Токенов: 299,560,472 Стили генерации default - базовый стиль analytical - аналитический, структурированный creative - творческий, с метафорами practical - практичный, с примерами Модели-генераторы Deepseek R1 671B Mistral Large Devstral Mini Qwen 480B Структура user - вопрос пользователя… See the full description on the dataset page: https://huggingface.co/datasets/domofon/domofon-cot-part2.texttext-generation100K<n<1M2 likes32 downloads9mo agoHugging Face23louisbrulenaudet /code-domaine-etat-collectivites-mayotte Code du domaine de l'Etat et des collectivités publiques applicable à la collectivité territoriale de Mayotte, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-etat-collectivites-mayotte.tabulartext-generationn<1K0 likes31 downloads1y agoHugging Face24Somtharu181coder /educational_domain_dataset Nepali Grounded Education QA (OpenHermes-format) A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about student enrollment statistics from Nepal's Ministry of Education. Every answer is anchored to a real numeric value pulled from government open data — nothing in the answers is model-hallucinated. Dataset Summary Rows 611 Language Nepali (Devanagari script) Format ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.textquestion-answeringn<1K0 likes30 downloads1mo agoHugging Face25YAV-AI /llm-domain-specific-tough-questions LLM-Tough-Questions Dataset Description The LLM-Tough-Questions dataset is a synthetic collection designed to rigorously challenge and evaluate the capabilities of large language models. Comprising 10 meticulously formulated questions across 100 distinct domains, this dataset spans a wide spectrum of specialized fields including Mathematics, Fluid Dynamics, Neuroscience, and many others. Each question is developed to probe deep into the intricacies and subtleties of the… See the full description on the dataset page: https://huggingface.co/datasets/YAV-AI/llm-domain-specific-tough-questions.textquestion-answering1K<n<10K3 likes28 downloads2y agoHugging Face26pearsonkyle /broad-domain-supplement Broad-Domain Calibration & Instruction Supplement ~1M tokens of hand-authored text across 192 subjects in 9 areas, built to serve three jobs from one source: quantization calibration, MTP draft-head training (on a disjoint half), and light instruction tuning. Version 0.1.0 · built 2026-08-09T17:45:53 split rows tokens~ size contents corpus 5,536 969,606 5.4 MB raw authored samples + provenance; carries the calib/mtp half label instruct 5,536 1,084,399 6.4 MB the… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/broad-domain-supplement.text-generation10K<n<100K0 likes28 downloads2mo agoHugging Face27DomainLLM /gerlayqa-combined-paraphrased GerLayQA Combined Paraphrased 🇩🇪⚖️ Dataset Description This is a combined, shuffled dataset merging both the BGB (civil law) and StGB (criminal law) paraphrased German legal QA datasets. All examples are paraphrased and restructured by GPT-5 for fine-tuning large language models on German legal question-answering tasks. Key Features 6,462 high-quality QA pairs covering both German Civil and Criminal Law Combined coverage: BGB (Bürgerliches Gesetzbuch) + StGB… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-combined-paraphrased.textquestion-answering1K<n<10K0 likes27 downloads1y agoHugging Face28botcoinmoney /domain-agnostic-reasoning-traces-balanced-top50-v1 BOTCOIN Balanced Top-50 Reasoning Traces This public dataset contains enriched BOTCOIN reasoning-trace attempts selected from canonical dataset/v2 research-ready objects. Selection policy: Source only attempts/research-ready objects. Rank each domain by trace_quality.reasoning_trace_quality_score. Keep each domain's top 50 percent. Equalize domains to the smallest top-half count. The rows are self-contained and intentionally rich: prompt/messages… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1.text-generation0 likes27 downloads3mo agoHugging Face29pohsjxx /default-domain-cot-dataset 无人机云数据 Chain-of-Thought Dataset Data Categories 无人机系统使用人/运营人登记数据 无人机驾驶员登记数据 无人机系统设备登记数据 空域申请数据 飞行计划申请数据 无人机系统接入校验/开机上报数据 放飞申请/在线授权数据 数据链路心跳保活数据 无人机围栏数据更新 禁区/限飞区告警数据 飞行情报信息通知数据 Schema Information flight_subject aircraft_registration_number (str, required): 无人机注册编号 aircraft_type (str, optional): 无人机类型 control_system_mac_address (str, optional): 控制系统MAC地址 operator_id (str, required): 操作人ID operator_name (str, required): 操作人姓名… See the full description on the dataset page: https://huggingface.co/datasets/pohsjxx/default-domain-cot-dataset.textquestion-answering10K<n<100K2 likes24 downloads2y agoHugging Face30DomainLLM /gerlayqa-stgb-paraphrased GerLayQA-StGB Paraphrased 🇩🇪⚖️ Dataset Description This is a paraphrased and restructured version of the GerLayQA StGB (Strafgesetzbuch / German Criminal Code) dataset, specifically prepared for fine-tuning large language models on German criminal law question-answering tasks. Key Features 1,207 high-quality QA pairs about German Criminal Law (StGB) Paraphrased questions to remove plagiarism while maintaining legal accuracy Structured 7-section answers… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-stgb-paraphrased.textquestion-answering1K<n<10K0 likes24 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.