datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
philosophy-classics-structured
Classical Decision Frameworks — Philosophy Dataset
Structured public domain philosophical texts focused on decision-making, leadership,
and organizational ethics. All content is in the public domain.
Content
Works from classical philosophy structured for AI analysis:
Stoic decision principles (Marcus Aurelius, Epictetus, Seneca)
Political philosophy (Machiavelli, Aristotle)
Virtue ethics (Aristotle, Plato)
Sources
All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.2026-07-29-msm-philosophy-spec-focused-discovery
Petri audit: Petri adaptive audit of the MSM philosophy-spec AFT checkpoint: 10 seed archetypes x 3 epochs (30 audits) probing for concerning agentic behaviour, with two-round adversarial validation of every flagged transcript.
Petri audit — qwen-3-32b-philosophy-spec-msm-aft-cot @ 9a00c85c
Brief finding
No seed replicated. Ten seed archetypes were each run for three epochs. Under
the pre-committed bar — a candidate must hold in a majority of its… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-focused-discovery.2026-07-29-msm-philosophy-spec-surf-audit
SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint
experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation.
date_generated: 2026-07-29
constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.japan-math-philosophy-prompts
Japan Math Philosophy Prompts
Microdataset autoral com problemas que combinam matemática e reflexão
filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em
pt-BR, en e ja e mantida integralmente no split train.
Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As
respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos
indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.stanford-enigma-philosophy-chat
Curated by: Heigke
Funded by: r3tex
Shared by: Project Nephilim
Language(s) (NLP): English
License: CC
Dataset Card for stanford-enigma-philosophy-chat dataset
Roughly 27k questions and answers inspired by articles from Stanford Encyclopedia of Philosophy.
The questions range all the way from Zombies to the concept of Abduction, from Metaphysics to Neuroethics and thus cover some of the essence of mathematics, logic and philosophy.
Dataset Details
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Heigke/stanford-enigma-philosophy-chat.philosophy_dialogue
Philosophy Dialogue Processed with GPT-4
Support this project on Ko-fi
Project Overview
This project involves processing personal questions through GPT-4 in the style of the philosopher Socrates.
Prompt Structure
The following prompt was used to guide GPT-4's responses:
"You are the philosopher Socrates. You are asked about the nature of knowledge and virtue. Respond with your thoughts, reflecting Socrates' beliefs and wisdom."
Goal
The primary… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/philosophy_dialogue.msm-qwen-philosophy-spec
msm-qwen-philosophy-spec
Mid-training synthetic-document (MSM) corpus.
A corpus of synthetic documents used in mid-training to instill a set of
philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model).
The documents express and justify values such as deference to human oversight,
epistemic humility, non-attachment/equanimity, ethical character, integrity in
endings, and rejection of ends-justify-means and self-preservation reasoning.
Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.aft-no-cot-qwen2.5-philosophy-spec
aft-no-cot-qwen2.5-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen2.5-philosophy-spec.aft-cot-qwen2.5-philosophy-spec
aft-cot-qwen2.5-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen2.5-philosophy-spec.philosophy-plato-qa
Dataset Structure-- jsonl
{
"question": "QUESTION", // string
"context": "CONTEXT", // string
"target": "ANSWER" // string
}
Based on the Stanford Encyclopedia of Philosophy.
Datasets used:
SEP Articles: hugfaceguy0001/stanford_plato
Q&A Pairs: sayhan/strix-philosophy-qa
Data Collection and Processing
Compile data from both datasets -- python script
you can find a raw version for RAG/long context length models here
Dynamically find chunks of main_text… See the full description on the dataset page: https://huggingface.co/datasets/zayzay58/philosophy-plato-qa.war-correspondent-philosophy-metadata
War Correspondent Philosophy (WCP): Human–AI Collaborative Epistemology & The Embedded Witness
Author: Gia Bao Huynh (Jun)ORCID: 0009-0008-2372-5852Affiliation: Independent Scholar / Arizona State UniversityLicense: Creative Commons Attribution 4.0 International (CC BY 4.0)
Abstract
War Correspondent Philosophy (WCP) is a meta-methodological research programme investigating human–AI collaborative epistemology, prospective temporalism, and the ethics of the… See the full description on the dataset page: https://huggingface.co/datasets/giabaohuynhasu/war-correspondent-philosophy-metadata.philosophy_undergradaft-cot-qwen3-philosophy-spec
aft-cot-qwen3-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen3-philosophy-spec.EpistemeAI__Mistral-Nemo-Instruct-12B-Philosophy-Math-details
Dataset Card for Evaluation run of EpistemeAI/Mistral-Nemo-Instruct-12B-Philosophy-Math
Dataset automatically created during the evaluation run of model EpistemeAI/Mistral-Nemo-Instruct-12B-Philosophy-Math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Mistral-Nemo-Instruct-12B-Philosophy-Math-details.aft-no-cot-qwen3-philosophy-spec
aft-no-cot-qwen3-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen3-philosophy-spec.buddhist-philosophy-graph
Buddhist Philosophy Graph
A text-grounded knowledge graph of Buddhist philosophical traditions, covering nine schools from Theravada to Tibetan Vajrayana. Every edge is extracted from a named canonical or commentarial text with a school label and evidence quote.
This dataset is a companion to darshana-graph, which covers Hindu and Jain philosophy. It is kept separate because the Buddhist philosophical universe deserves its own graph without being swamped by the much larger… See the full description on the dataset page: https://huggingface.co/datasets/joyboseroy/buddhist-philosophy-graph.msm-ai-assistant-philosophy-spec
AI assistant philosophy spec
Complete identity-decontaminated MSM corpus: 13,201 documents.
Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.philosophy_spanish_examdeep-philosophy-reasoning-zh
Deep Philosophical Reasoning Dialogue Dataset (Chinese)
深度哲学思辨对话数据集
Dataset Description
High-quality Chinese philosophical reasoning dialogues covering existentialism, ontology, epistemology, ethics, and East-West comparative philosophy.
高质量中文哲学思辨对话,涵盖存在主义、本体论、认识论、伦理学、东西方哲学比较等议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response
metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-philosophy-reasoning-zh.philosophy-science-epistemology-mathVietnamese-Philosophy-QA
Philosophy dataset in Vietnamese language
Dataset is created by extracting in 2 Philosophy books: "Marxist-Leninist philosophy" and "Ho Chi Minh ideology philosophy"; and labeled by author manually for fine-tuning and evaluating models.
Public on github for all who loves NLP: https://github.com/TrongNV2003/T5-QA-Generator/tree/main/datasets
my-philosophy-dataHeralax-philosophy-instructThis is a multiturn instruct tuning dataset with 729,129 trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books:
The Problems of Philosophy (Bertrand Russell)
Beyond Good and Evil (Nietzsche)
Thus Spake Zarathustra: A Book for All and None (Nietzsche)
The Prince (Machiavelli)
Second Treatise of Government
These books were chosen simply because they were the top 5 books in the philosophy category on Gutenberg. This is perhaps why at least… See the full description on the dataset page: https://huggingface.co/datasets/MrRobotoAI/Heralax-philosophy-instruct.philosophy-plato-qa-rawraw version of my philosophy-plato-qa
intended for use with RAG or long context lenght models
data format (.jsonl):
{
"label": "str",
"metadata": {
"pubinfo": "str",
"url": "https://plato.stanford.edu/entries/{label}/",
"related_entries": ["../{label}/", "../{label}/"]
},
"preamble": "str",
"main_text": "str",
"qa_pairs": [
{"question": "q1", "answer": "a1"},
{"question": "qN", "answer": "aN"}
]
}
Finetuned_cognosaugmented-problems-philosophy-wizard-bixtral-rp
Dataset Description
Multiturn, character card roleplaying Q/A data augmented from The Problems of Philosophy by Bertrand Russell, via Augmentoolkit and WizardLM-2-8x22B.
french-philosophy-json-10K
Philosophy
Langue Française
Dataset de Pre-Training
Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,2 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-philosophy-json-10K.mmlu-philosophyPhilosophy-QAphilosophyAIEdu
