datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.blobfish-domainbench-24
Blobfish DomainBench-24 v3.3.2
DomainBench-24 v3.3.2 is the public release record for 24 realistic stateful agent tasks across six professional domains. Every task runs in a checked-in SQLite-backed MCP world with typed read/write tools and deterministic state, exact argument-aware trace, containment, persisted causal workpaper, stakeholder handoff, provider-native exact-record readback for every changed domain row, and persisted-result readback verification.
The release… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24.Open-domain_Financial_QA
Large-scale Open-domain Financial QA (LOFin)
This repository accompanies the dataset used in the paper:
Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents
(ACL 2025 Findings)
We introduce a benchmark for open-domain financial question answering over standardized documents, with a focus on multi-document and multi-hop reasoning.
📊 Benchmark Overview
This benchmark targets realistic financial QA tasks… See the full description on the dataset page: https://huggingface.co/datasets/HYdsl/Open-domain_Financial_QA.Domofon-Cot-Conversations-700k
Domofon-Cot-Conversations-700k
Synthetic XML conversation data for training small language models on reasoning,
instruction following, XML formatting, and tool-use traces.
Repository: domofon/Domofon-Cot-Conversations-700k
What is inside
The dataset contains cleaned generated XML conversations from six families:
conv: multi-turn factual conversations with tool-use traces.
instruct: text-processing instructions, including deterministic count tool calls.
ds:… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Domofon-Cot-Conversations-700k.domain-agnostic-causal-reasoning-tuning
Domain-Agnostic Causal Reasoning Tuning Dataset
Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network.
The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.Dendrite-Synth-Multi-Domain
Dendrite Synth Multi-Domain
A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer
triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.domofon-cot-part1
Domofon COT Dataset - Part 1
Mixed Chain-of-Thought датасет на русском языке. (Russian language)
Статистика
Строк | lines : 700,000
Токенов: | tokens 594,412,523
Модели-генераторы (Distill Generation Models used):
Deepseek R1 671B
Mistral Large
Devstral Mini
Qwen 480B
Структура (Structure)
user - вопрос пользователя (User Question)
thinking - цепочка рассуждений (Chain-of-Thought)
answer - финальный ответ (Final answer)
Связанные… See the full description on the dataset page: https://huggingface.co/datasets/domofon/domofon-cot-part1.domains
Domains dataset
Documentation coming soon
Dendrite-Synth-Multi-Domain
Dendrite Synth Multi-Domain
A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer
triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.omnimcp_browser_dom_structured_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.DomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset]
This Dataset includes 980,065 geographic names as of September 10, 2023.
It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories.
Example: feature_name: Abercrombie Gulch
GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.Mid-Training_data_of_separate_domains
Breaking the Data Barrier – Building GUI Agents Through Task Generalization
This is the official dataset repository of GUIMid
1. Data Overview
AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks.
The performances of different domains as mid-training data are as follows:
Domains
Observation
WebArena (PR)
WebArena (SR)
AndroidWorld (SR)
GUI Post-Training Only
Image
26.3
6.2
9.0
Public Baselines
GPT-4o-2024-11-20
Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.ft-llm-2026-domain-specific-qa
FT-LLM 2026 Domain-Specific QA
A Japanese financial-domain visual QA dataset used for Phase 3 domain fine-tuning of the COMPASS Vision-Language Model. Question–answer pairs were generated with Qwen3-VL from scraped Japanese government financial PDFs (Cabinet Office, Financial Services Agency, Ministry of Finance), covering four difficulty tiers: (A) numeric extraction, (B) rate-of-change & comparison, (C) financial formula application, and (D) complex reasoning. Each answer includes… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-domain-specific-qa.tariff_trade_domain.synthetic_trade_qa_kr
Korean Trade Domain QA Dataset
A Korean-language question-answering dataset for the international trade domain, combining 21,399 QA pairs from three complementary sources: official 무역영어 1급 certification exam questions, trade terminology definitions, and lecture-derived QA pairs. Designed for fine-tuning and evaluating LLMs on Korean trade domain knowledge.
Dataset Description
The dataset covers vocabulary, regulations, procedures, and concepts from Korean… See the full description on the dataset page: https://huggingface.co/datasets/lablup/tariff_trade_domain.synthetic_trade_qa_kr.Multi-Domain-Reasoning-SFT
Multi-Domain-Reasoning-SFT
Dataset Summary
The Multi-Domain-Reasoning-SFT dataset by EnDevSols is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset designed to train large language models in deep reasoning, technical analysis, and complex problem-solving.
Consisting of nearly 580,000 meticulously structured examples, this dataset is specifically engineered to teach models how to "think" before they answer. It separates the internal cognitive process… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/Multi-Domain-Reasoning-SFT.Legal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.science_behavioral_and_domain_diversity_dataset
Nepali Science SFT Dataset — Clean Candidate
A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script.
This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering.
Dataset Overview
Property
Value
Dataset file
clean_candidate.jsonl
Records
29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.Code-170k-dombe
Dataset Description
Code-170k-dombe is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Dombe, making coding education accessible to Dombe speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Dombe language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-dombe.Domain-ViWikiQA
Vietnamese Domain WikiQA Dataset
Dataset Overview
Vietnamese WikiQA is a Vietnamese question-answering dataset built from Vietnamese Wikipedia. It is designed for QA research, dataset analysis, and difficulty modeling.
The dataset contains 7,592 QA pairs after generation, automatic filtering, human verification, and release normalization.
Dataset Structure
Each row in the public release represents one QA instance generated from a Wikipedia section… See the full description on the dataset page: https://huggingface.co/datasets/HienNGuit/Domain-ViWikiQA.gerlayqa-bgb-paraphrased
GerLayQA-BGB Paraphrased 🇩🇪⚖️
Dataset Description
This is a paraphrased and restructured version of the GerLayQA BGB (Bürgerliches Gesetzbuch / German Civil Code) dataset, specifically prepared for fine-tuning large language models on German civil law question-answering tasks.
Key Features
5,255 high-quality QA pairs about German Civil Law (BGB)
Paraphrased questions to remove plagiarism while maintaining legal accuracy
Structured 7-section answers following… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-bgb-paraphrased.domofon-cot-part2
Domofon COT Dataset - Part 2 (Mixed Styles)
Mixed Chain-of-Thought датасет на русском языке с разными стилями генерации.
Статистика
Строк: 321,072
Токенов: 299,560,472
Стили генерации
default - базовый стиль
analytical - аналитический, структурированный
creative - творческий, с метафорами
practical - практичный, с примерами
Модели-генераторы
Deepseek R1 671B
Mistral Large
Devstral Mini
Qwen 480B
Структура
user - вопрос пользователя… See the full description on the dataset page: https://huggingface.co/datasets/domofon/domofon-cot-part2.code-domaine-etat-collectivites-mayotte
Code du domaine de l'Etat et des collectivités publiques applicable à la collectivité territoriale de Mayotte, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-etat-collectivites-mayotte.educational_domain_dataset
Nepali Grounded Education QA (OpenHermes-format)
A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about
student enrollment statistics from Nepal's Ministry of Education. Every answer is
anchored to a real numeric value pulled from government open data — nothing in the
answers is model-hallucinated.
Dataset Summary
Rows
611
Language
Nepali (Devanagari script)
Format
ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.llm-domain-specific-tough-questions
LLM-Tough-Questions Dataset
Description
The LLM-Tough-Questions dataset is a synthetic collection designed to rigorously challenge and evaluate the capabilities of large language models. Comprising 10 meticulously formulated questions across 100 distinct domains, this dataset spans a wide spectrum of specialized fields including Mathematics, Fluid Dynamics, Neuroscience, and many others. Each question is developed to probe deep into the intricacies and subtleties of the… See the full description on the dataset page: https://huggingface.co/datasets/YAV-AI/llm-domain-specific-tough-questions.broad-domain-supplement
Broad-Domain Calibration & Instruction Supplement
~1M tokens of hand-authored text across 192 subjects in 9 areas, built to serve three jobs from one source: quantization calibration, MTP draft-head training (on a disjoint half), and light instruction tuning.
Version 0.1.0 · built 2026-08-09T17:45:53
split
rows
tokens~
size
contents
corpus
5,536
969,606
5.4 MB
raw authored samples + provenance; carries the calib/mtp half label
instruct
5,536
1,084,399
6.4 MB
the… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/broad-domain-supplement.gerlayqa-combined-paraphrased
GerLayQA Combined Paraphrased 🇩🇪⚖️
Dataset Description
This is a combined, shuffled dataset merging both the BGB (civil law) and StGB (criminal law) paraphrased German legal QA datasets. All examples are paraphrased and restructured by GPT-5 for fine-tuning large language models on German legal question-answering tasks.
Key Features
6,462 high-quality QA pairs covering both German Civil and Criminal Law
Combined coverage: BGB (Bürgerliches Gesetzbuch) + StGB… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-combined-paraphrased.domain-agnostic-reasoning-traces-balanced-top50-v1
BOTCOIN Balanced Top-50 Reasoning Traces
This public dataset contains enriched BOTCOIN reasoning-trace attempts selected
from canonical dataset/v2 research-ready objects.
Selection policy:
Source only attempts/research-ready objects.
Rank each domain by trace_quality.reasoning_trace_quality_score.
Keep each domain's top 50 percent.
Equalize domains to the smallest top-half count.
The rows are self-contained and intentionally rich: prompt/messages… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1.default-domain-cot-dataset
无人机云数据 Chain-of-Thought Dataset
Data Categories
无人机系统使用人/运营人登记数据
无人机驾驶员登记数据
无人机系统设备登记数据
空域申请数据
飞行计划申请数据
无人机系统接入校验/开机上报数据
放飞申请/在线授权数据
数据链路心跳保活数据
无人机围栏数据更新
禁区/限飞区告警数据
飞行情报信息通知数据
Schema Information
flight_subject
aircraft_registration_number (str, required): 无人机注册编号
aircraft_type (str, optional): 无人机类型
control_system_mac_address (str, optional): 控制系统MAC地址
operator_id (str, required): 操作人ID
operator_name (str, required): 操作人姓名… See the full description on the dataset page: https://huggingface.co/datasets/pohsjxx/default-domain-cot-dataset.gerlayqa-stgb-paraphrased
GerLayQA-StGB Paraphrased 🇩🇪⚖️
Dataset Description
This is a paraphrased and restructured version of the GerLayQA StGB (Strafgesetzbuch / German Criminal Code) dataset, specifically prepared for fine-tuning large language models on German criminal law question-answering tasks.
Key Features
1,207 high-quality QA pairs about German Criminal Law (StGB)
Paraphrased questions to remove plagiarism while maintaining legal accuracy
Structured 7-section answers… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-stgb-paraphrased.
