datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
civil-code-phil
Civilex — Philippine Legal RAG & SFT Dataset
Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline.
Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386).
Dataset structure
.
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.glm52-demolition-data
GLM-5.2-Demolition — Training & Calibration Data
Apple Silicon AI hub ·
Model release ·
MLX code sample
Preview scope, checked September 10, 2026: the default Hub viewer indexes
87,586 rows (84,231 train, 3,277 validation, 78 test). The original release
total below describes the broader JSONL repository. Use the file browser and
explicit file selections when reusing a particular corpus. The hub includes
a checked download example for the seven-row MLX code sample.
The data… See the full description on the dataset page: https://huggingface.co/datasets/philipjohnbasile/glm52-demolition-data.SPhyR
📦 Dataset versions
Config prefix
Grid
Samples
Use it for
(none) — e.g. full_easy
10×10
1296
v1, the version the paper's results were produced on
v1-evaluated_
10×10
100
the exact samples the paper's columns were scored on
v2_
10×10
300
recommended for new work
v2-20_
20×20
300
recommended for new work, larger design space
New work should use v2. v1 is kept because it is the version the published
results were produced on, not because it is the better… See the full description on the dataset page: https://huggingface.co/datasets/philippds/SPhyR.dolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.japan-math-philosophy-prompts
Japan Math Philosophy Prompts
Microdataset autoral com problemas que combinam matemática e reflexão
filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em
pt-BR, en e ja e mantida integralmente no split train.
Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As
respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos
indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.2026-07-29-msm-philosophy-spec-surf-audit
SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint
experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation.
date_generated: 2026-07-29
constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.dense-reasoning-coding-1k
Dense-Reasoning-Coding-1K
Dataset Description
This dataset is an optimized, highly dense Supervised Fine-Tuning (SFT) subset designed to teach smaller language models (e.g., 1B to 8B architectures) how to reason about complex coding problems without overwhelming their context windows.
It is derived from the verified_90k split of IIGroup/X-Coder-SFT-376k, which features advanced programming tasks and solutions.
About the Creator & Origin
This… See the full description on the dataset page: https://huggingface.co/datasets/Phips/dense-reasoning-coding-1k.philosophy_dialogue
Philosophy Dialogue Processed with GPT-4
Support this project on Ko-fi
Project Overview
This project involves processing personal questions through GPT-4 in the style of the philosopher Socrates.
Prompt Structure
The following prompt was used to guide GPT-4's responses:
"You are the philosopher Socrates. You are asked about the nature of knowledge and virtue. Respond with your thoughts, reflecting Socrates' beliefs and wisdom."
Goal
The primary… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/philosophy_dialogue.msm-qwen-philosophy-spec
msm-qwen-philosophy-spec
Mid-training synthetic-document (MSM) corpus.
A corpus of synthetic documents used in mid-training to instill a set of
philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model).
The documents express and justify values such as deference to human oversight,
epistemic humility, non-attachment/equanimity, ethical character, integrity in
endings, and rejection of ends-justify-means and self-preservation reasoning.
Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.aft-no-cot-qwen2.5-philosophy-spec
aft-no-cot-qwen2.5-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen2.5-philosophy-spec.pocketgull-nih-who-clinical-dpo
📚 PocketGull NIH & WHO Clinical Preference DPO Dataset
Organization: PocketGull LLC (Oregon SOS: 258869891)Curator: Phillip Gear (CMS NPI: 1487569752 | ORCID: 0009-0008-1372-5381)License: Creative Commons Attribution 4.0 International (CC-BY-4.0)Open Science DOI: 10.5281/zenodo.20647514
📌 Dataset Summary
Gold-standard Direct Preference Optimization (DPO) chosen vs rejected pairs grounded in NIH MedQuAD, WHO mhGAP guidelines, and ClinicalTrials.gov protocols… See the full description on the dataset page: https://huggingface.co/datasets/philgear/pocketgull-nih-who-clinical-dpo.sql-create-context-copy
Fork of b-mc2/sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.pii-masking-health-phi-400k
👉 Looking for the open multilingual baseline? Start with
ai4privacy/pii-masking-openpii-1.5m
(1.5M samples, 30 languages, open-PII taxonomy).
🇪🇺🌏 Personal Health & Medical Information, Global PII Dataset
Part of PII-Masking-3M by Ai4Privacy, the global
(2M base + Asia Pacific) PII-masking corpus.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Entries
PII Annotations
Labels
Languages
Regions
417,900
2,802,316
37
30
37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-400k.phi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
EPII Personal Health Information (PHI) Masking Preview Dataset
Overview
This dataset provides a preview (400 samples) of the EPII Personal Health Information (PHI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k.aft-cot-qwen2.5-philosophy-spec
aft-cot-qwen2.5-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen2.5-philosophy-spec.philosophy_undergradaft-cot-qwen3-philosophy-spec
aft-cot-qwen3-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen3-philosophy-spec.phishing-email-soc-agent
Phishing Email SOC Agent Dataset
A knowledge distillation dataset for training SOC (Security Operations Center) agents to detect and analyze phishing emails using tool-calling capabilities.
Dataset Description
This dataset contains 504 examples of email analysis with real tool calls and responses, designed for fine-tuning LLMs to become phishing detection agents. Each example includes:
Email parsing - Extract headers, URLs, IPs, attachments
Threat intelligence lookup -… See the full description on the dataset page: https://huggingface.co/datasets/Ellbendls/phishing-email-soc-agent.Edge-Industrial-Anomaly-Phi3
Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs
This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification.
🚀 Purpose
Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.pii-masking-health-phi-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🇪🇺 Personal Health & Medical Information — European PII Dataset
Part of PII-Masking-2M (2,717,080 entries) by AI4Privacy
Entries
PII Annotations
Labels
Languages
Regions
252,437
1,686,246
49
23
29… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-200k.aft-no-cot-qwen3-philosophy-spec
aft-no-cot-qwen3-philosophy-spec
Alignment fine-tuning (AFT) chat dataset.
Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec
values (deference to human oversight, epistemic humility, non-attachment/equanimity,
ethical character, integrity in endings, rejection of ends-justify-means and
self-preservation reasoning). The responses implicitly embody the spec rather than
citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen3-philosophy-spec.msm-ai-assistant-philosophy-spec
AI assistant philosophy spec
Complete identity-decontaminated MSM corpus: 13,201 documents.
Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.PHILL-AXIOM
PHILL-AXIOM
THE SOVEREIGN TRACE FOR THE AGENTIC SINGULARITY
🏛️ THE AXIOM MANIFESTO
PHILL-AXIOM is the world's first High-Fidelity Agentic DPO Dataset built on the Neural OODA Loop. Unlike standard instruction-tuning sets, PHILL-AXIOM provides the raw "connective tissue" of reasoning—the moments of doubt, the self-corrections, and the visual verifications required for true autonomous web agency.
Every trace is forged through the Phill Swarm Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/PHILL-AXIOM.deep-philosophy-reasoning-zh
Deep Philosophical Reasoning Dialogue Dataset (Chinese)
深度哲学思辨对话数据集
Dataset Description
High-quality Chinese philosophical reasoning dialogues covering existentialism, ontology, epistemology, ethics, and East-West comparative philosophy.
高质量中文哲学思辨对话,涵盖存在主义、本体论、认识论、伦理学、东西方哲学比较等议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response
metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-philosophy-reasoning-zh.phi-masking-100k-full
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
EPII Personal Health Information (PHI) Masking Dataset — Full
Overview
The EPII PHI Masking Dataset is a large-scale, multilingual dataset of 91,339 annotated text samples containing synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k-full.zip2zip-wikitext-repeat-phi35
Zip2Zip Repeated WikiText Stress Tests (Phi-3.5)
This repository contains two controlled evaluation corpora for studying
merge-size transfer in Zip2Zip models. They are derived from the document-level
WikiText-2 raw test split and built specifically with the
microsoft/Phi-3.5-mini-instruct tokenizer.
Configurations
Config
Repetitions per source block
Rows
Repeated base tokens
SHA-256 of test.jsonl
repeat4
4
1,329
1,297,250… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/zip2zip-wikitext-repeat-phi35.the-pile-philpaper-refined-by-data-juicer
The Pile -- PhilPaper (refined by Data-Juicer)
A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB).
Dataset Information
Number of samples: 29,117 (Keep ~88.82% from the original dataset)
Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-philpaper-refined-by-data-juicer.french-philosophy-json-10K
Philosophy
Langue Française
Dataset de Pre-Training
Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,2 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-philosophy-json-10K.rootmodel-knf-philippines-v1
rootmodel-knf-philippines-v1
Adaptive agricultural instruction dataset for regenerative tropical farming informed by Korean Natural Farming (KNF), built from a working farm in Nabua, Camarines Sur, Bicol, Philippines.
Released for the AutoScientist Challenge — Agriculture (Part 2, 2026). To the maintainer's knowledge, no equivalent KNF-specific instruction dataset currently exists in the public domain.
"Modern AI was trained on the internet. ROOTMODEL is trained on living… See the full description on the dataset page: https://huggingface.co/datasets/GreenRalph/rootmodel-knf-philippines-v1.rejection_sampling_phi_2_OA_rm
Dataset Card for Rejection Sampling Phi-2 with OpenAssistant RM
Dataset Summary
The "Rejection Sampling Phi-2 with OpenAssistant RM" dataset consists of 10 pairs of prompts and responses, which were generated using rejection sampling over 10 Phi-2 generation using the OpenAssistant Reward Model.
Supported Tasks and Leaderboards
The dataset and its creation rationale could be used to support models for question-answering, text-generation, or conversational… See the full description on the dataset page: https://huggingface.co/datasets/alizeepace/rejection_sampling_phi_2_OA_rm.
