philosophy
OpenHermes-2.5-Strix-Philosophy-Mistral-7B-LoRAMistral-Nemo-Instruct-12B-Philosophy-Math-i1-GGUFLlama3-stanford-encyclopedia-philosophy-QA-i1-GGUFt0-mt-3b-base-philosophy-sdfruggsea_-_Llama3-stanford-encyclopedia-philosophy-QA-ggufMistral-Nemo-Instruct-12B-Philosophy-Math-GGUFOpenHermes-2.5-Strix-Philosophy-Mistral-7B-LoRA-i1-GGUFphilosophy-mistral-i1-GGUF
stanford-encyclopedia-philosophy
Stanford Encyclopedia Philosophy (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/stanford-encyclopedia-philosophy", split = 'train')
philosophy-corpus
Philosophy & Humanities Corpus
Combined humanities and Wikipedia corpus for training small language models.
Dataset
Split
Lines
Size
Description
train.txt
3.0M
549 MB
Humanities (368K lines) + WikiText-103 (2.6M lines)
val.txt
315K
57 MB
Matching validation split
Sources
Humanities (368K lines, 66 MB)
54 classical philosophy and humanities texts:
Category
Works
Plato
Republic, Apology, Symposium, Phaedo, Crito, Meno… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/philosophy-corpus.The-Philosophy-Data-Project
About dataset
The Philosophy Data Project is a corpus and a set of anaylsis based philosophy texts, totaling over 50 texts and 30 authors, made by Kourosh Alizadeh.
school: Broad categorization of which school of thought each book belongs to. Sometimes, this classification can be vague or depend on interpretation. Thankfully, texts in this corpus are all distinctive examples of respective school of thought, so at leat here they are reasonable.
sentence_spacy and sentence_str:… See the full description on the dataset page: https://huggingface.co/datasets/yjkim27/The-Philosophy-Data-Project.2026-07-29-msm-philosophy-spec-petri-validation
Petri raw transcripts and validation: pilot, focused discovery, C5b control, rate estimation
experiment: The complete raw Petri (Inspect) audit corpus for the MSM out-of-distribution vulnerability investigation - every audit phase from the failed 4-audit pilot through the 30-audit focused discovery, the C5b control, the 3-seed/8-10-epoch rate-estimation re-run, and small Claude-subscription-auditor architecture trials - plus every validation artifact derived from them… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-petri-validation.Philosophy-Setsstanford-encyclopedia-of-philosophy_instruct
Description
This is a semi-synthetic instruct dataset meant for supervised finetuning of a large language model for the task of answering philosophical questions in a formal manner. The dataset is based on the Stanford Encyclopedia of Philosophy (SEP). Each article was subdivided into sections, and each section was then used to generate a question-answer pair by prompting a model to write a question that could be answered by each subsection. Subsection with a too high (>2000) or too… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/stanford-encyclopedia-of-philosophy_instruct.
