CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AiresPucrs /stanford-encyclopedia-philosophy Stanford Encyclopedia Philosophy (Teeny-Tiny Castle) This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research. How to Use from datasets import load_dataset dataset = load_dataset("AiresPucrs/stanford-encyclopedia-philosophy", split = 'train') texttext-classification100K<n<1M53 likes751 downloads2y agoHugging Face02LisaMegaWatts /philosophy-corpus Philosophy & Humanities Corpus Combined humanities and Wikipedia corpus for training small language models. Dataset Split Lines Size Description train.txt 3.0M 549 MB Humanities (368K lines) + WikiText-103 (2.6M lines) val.txt 315K 57 MB Matching validation split Sources Humanities (368K lines, 66 MB) 54 classical philosophy and humanities texts: Category Works Plato Republic, Apology, Symposium, Phaedo, Crito, Meno… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/philosophy-corpus.texttext-generation10M<n<100M0 likes323 downloads7mo agoHugging Face03ruggsea /stanford-encyclopedia-of-philosophy_instruct Description This is a semi-synthetic instruct dataset meant for supervised finetuning of a large language model for the task of answering philosophical questions in a formal manner. The dataset is based on the Stanford Encyclopedia of Philosophy (SEP). Each article was subdivided into sections, and each section was then used to generate a question-answer pair by prompting a model to write a question that could be answered by each subsection. Subsection with a too high (>2000) or too… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/stanford-encyclopedia-of-philosophy_instruct.texttext-generation10K<n<100K17 likes188 downloads5mo agoHugging Face04guicybercode /japan-math-philosophy-prompts Japan Math Philosophy Prompts Microdataset autoral com problemas que combinam matemática e reflexão filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em pt-BR, en e ja e mantida integralmente no split train. Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.textquestion-answeringn<1K0 likes119 downloads29d agoHugging Face05dougalldeepmind /2026-07-29-msm-philosophy-spec-surf-audit SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation. date_generated: 2026-07-29 constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.texttext-generationn<1K0 likes115 downloads2mo agoHugging Face06Hypersniper /philosophy_dialogue Philosophy Dialogue Processed with GPT-4 Support this project on Ko-fi Project Overview This project involves processing personal questions through GPT-4 in the style of the philosopher Socrates. Prompt Structure The following prompt was used to guide GPT-4's responses: "You are the philosopher Socrates. You are asked about the nature of knowledge and virtue. Respond with your thoughts, reflecting Socrates' beliefs and wisdom." Goal The primary… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/philosophy_dialogue.texttext-generationn<1K15 likes88 downloads3y agoHugging Face07Lots-of-LoRAs /task726_mmmlu_answer_generation_philosophy Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task726_mmmlu_answer_generation_philosophy Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task726_mmmlu_answer_generation_philosophy.texttext-generationn<1K0 likes80 downloads2y agoHugging Face08chloeli /msm-qwen-philosophy-spec msm-qwen-philosophy-spec Mid-training synthetic-document (MSM) corpus. A corpus of synthetic documents used in mid-training to instill a set of philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model). The documents express and justify values such as deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, and rejection of ends-justify-means and self-preservation reasoning. Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.texttext-generation10K<n<100K0 likes77 downloads4mo agoHugging Face09chloeli /aft-no-cot-qwen2.5-philosophy-spec aft-no-cot-qwen2.5-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen2.5-philosophy-spec.texttext-generation1K<n<10K0 likes70 downloads4mo agoHugging Face10ruggsea /stanford-encyclopedia-of-philosophy_chat_multi_turn_mistral_largeThis dataset is essentially identical to the Stanford Encyclopedia of Philosophy Chat Multi-turn Dataset, with one key difference: it uses Mistral Large 2 for conversation generation instead of LLaMA 3.1 70B. All other aspects, including format, statistics, and intended use, remain the same as the original dataset. texttext-generation10K<n<100K5 likes65 downloads2y agoHugging Face11chloeli /aft-cot-qwen2.5-philosophy-spec aft-cot-qwen2.5-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen2.5-philosophy-spec.texttext-generation1K<n<10K0 likes52 downloads4mo agoHugging Face12mstyslavity /philosophy_undergradtexttext-generation100K<n<1M0 likes50 downloads7mo agoHugging Face13chloeli /aft-cot-qwen3-philosophy-spec aft-cot-qwen3-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen3-philosophy-spec.texttext-generation1K<n<10K0 likes47 downloads4mo agoHugging Face14ruggsea /stanford-encyclopedia-of-philosophy_chat_multi_turn Multi-turn Stanford Encyclopedia of Philosophy Chat Dataset This dataset is designed for fine-tuning large language models to engage in multi-turn philosophical discussions while adopting the persona of a Philosophy professor named Phil. The resulting model should be able to converse like a university-level philosophy professor, who excels at explanations. This is a semi-synthetic dataset based on the Stanford Encyclopedia of Philosophy (SEP). It simulates conversations between Phil… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/stanford-encyclopedia-of-philosophy_chat_multi_turn.texttext-generation10K<n<100K14 likes43 downloads2y agoHugging Face15chloeli /aft-no-cot-qwen3-philosophy-spec aft-no-cot-qwen3-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen3-philosophy-spec.texttext-generation1K<n<10K0 likes34 downloads4mo agoHugging Face16ruggsea /stanford-encyclopedia-of-philosophy_chat_multi_turn_atheneThis dataset is essentially identical to the Stanford Encyclopedia of Philosophy Chat Multi-turn Dataset, with one key difference: it uses Athene 70B for conversation generation instead of LLaMA 3.1 70B. All other aspects, including format, statistics, and intended use, remain the same as the original dataset. texttext-generation10K<n<100K2 likes33 downloads2y agoHugging Face17P0u4a /msm-ai-assistant-philosophy-spec AI assistant philosophy spec Complete identity-decontaminated MSM corpus: 13,201 documents. Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.texttext-generation10K<n<100K0 likes33 downloads13d agoHugging Face18AI-Culture-Commons /philosophy-culture-translations-html-csv AI-Culture Philosophy and Culture Translations CSV + HTML Corpus The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind. This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.imagetranslation1K<n<10K2 likes30 downloads1y agoHugging Face19AngelWarmSmile123 /deep-philosophy-reasoning-zh Deep Philosophical Reasoning Dialogue Dataset (Chinese) 深度哲学思辨对话数据集 Dataset Description High-quality Chinese philosophical reasoning dialogues covering existentialism, ontology, epistemology, ethics, and East-West comparative philosophy. 高质量中文哲学思辨对话,涵盖存在主义、本体论、认识论、伦理学、东西方哲学比较等议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output: AI response metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-philosophy-reasoning-zh.texttext-generation1K<n<10K1 likes27 downloads3mo agoHugging Face20Dorian2B /french-philosophy-10K Philosophy Langue Française Dataset de Pre-Training Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,2 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning de… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-philosophy-10K.texttext-generation10K<n<100K2 likes22 downloads1y agoHugging Face21asheinin /The_Mathematical_Principles_of_Natural_Philosophy_1846 Dataset Card for Newton Matematical Principles Dataset Summary This dataset is meant to me used as a showcase for finetuning an LLM on a specific domain. Supported Tasks and Leaderboards Text generation Languages English Dataset Structure text file Data Splits Train only, the entire 1846 English version of the book. Source Data… See the full description on the dataset page: https://huggingface.co/datasets/asheinin/The_Mathematical_Principles_of_Natural_Philosophy_1846.texttext-generation1K<n<10K3 likes16 downloads3y agoHugging Face22Dorian2B /french-philosophy-json-10K Philosophy Langue Française Dataset de Pre-Training Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,2 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-philosophy-json-10K.texttext-generation10K<n<100K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.