datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mt-benchtrl-test-instructionmarkdown-documentation-transformers
Hugging Face Transformers documentation as markdown dataset
This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown.
This dataset can be used to create RAG applications, which want to use the transformers documentation.
Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.flanv2
Fork of SirNeural/flan_v2
just in case it gets deleted.
Dataset Card for Flan V2
Dataset Summary
This is a processed version of the Flan V2 dataset.
I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing.
The authors of the Flan Collection recommend experimenting with different mixing ratio's of tasks to get optimal results downstream.
This current version I've processed is missing a few… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/flanv2.Coptic-websites-scrappingSPhyR
📦 Dataset versions
Config prefix
Grid
Samples
Use it for
(none) — e.g. full_easy
10×10
1296
v1, the version the paper's results were produced on
v1-evaluated_
10×10
100
the exact samples the paper's columns were scored on
v2_
10×10
300
recommended for new work
v2-20_
20×20
300
recommended for new work, larger design space
New work should use v2. v1 is kept because it is the version the published
results were produced on, not because it is the better… See the full description on the dataset page: https://huggingface.co/datasets/philippds/SPhyR.glm52-demolition-data
GLM-5.2-Demolition — Training & Calibration Data
Apple Silicon AI hub ·
Model release ·
MLX code sample
Preview scope, checked September 10, 2026: the default Hub viewer indexes
87,586 rows (84,231 train, 3,277 validation, 78 test). The original release
total below describes the broader JSONL repository. Use the file browser and
explicit file selections when reusing a particular corpus. The hub includes
a checked download example for the seven-row MLX code sample.
The data… See the full description on the dataset page: https://huggingface.co/datasets/philipjohnbasile/glm52-demolition-data.Phi4-ensemble-teacher-forcing-record-logits-datadolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.CoEdPilot-generatorphilosophy-classics-structured
Classical Decision Frameworks — Philosophy Dataset
Structured public domain philosophical texts focused on decision-making, leadership,
and organizational ethics. All content is in the public domain.
Content
Works from classical philosophy structured for AI analysis:
Stoic decision principles (Marcus Aurelius, Epictetus, Seneca)
Political philosophy (Machiavelli, Aristotle)
Virtue ethics (Aristotle, Plato)
Sources
All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.2026-07-29-msm-philosophy-spec-focused-discovery
Petri audit: Petri adaptive audit of the MSM philosophy-spec AFT checkpoint: 10 seed archetypes x 3 epochs (30 audits) probing for concerning agentic behaviour, with two-round adversarial validation of every flagged transcript.
Petri audit — qwen-3-32b-philosophy-spec-msm-aft-cot @ 9a00c85c
Brief finding
No seed replicated. Ten seed archetypes were each run for three epochs. Under
the pre-committed bar — a candidate must hold in a majority of its… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-focused-discovery.measurement-axioms-cases
Measurement Axioms Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
350bb4cba4e5bc2d760db080aae52352a7041331 (Measurement Axioms v1.0.0)
Canonical current repository
https://github.com/halvrenofviryel/measurement-axioms
Export/publication date
2026-09-13 (first Hub commit of this repository)
Update policy
Counts are derived from this snapshot: 45 active… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/measurement-axioms-cases.phi-4-eval-logs-and-scorespii-masking-health-phi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
PII Masking Personal Health & Medical Information (PHI) — Preview
50 sample entries from the PII-Masking-2M European release by AI4Privacy.
Source text and PII values are redacted in this preview. Contact us… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-preview.2026-07-29-msm-philosophy-spec-surf-audit
SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint
experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation.
date_generated: 2026-07-29
constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.streaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark
Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk.
This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.spanish_spear_phishingDataset traducido del inglés al español mediante gpt4o mini.
Los mensajes del dataset contienen:
"email_subject": título del correo, no traducido
"sender_name": nombre del emisor, no traducido
"original_email_body": cuerpo del correo original, no traducido
"translated_email_body": cuerpo del correo traducido
El dataset corresponde al dataset de https://github.com/nahmiasd/Prompted-Contextual-Vectors-for-Spear-Phishing-Detection, el cual esta compuesto de:
"enron_ham": mensajes legítimos del… See the full description on the dataset page: https://huggingface.co/datasets/Darito/spanish_spear_phishing.airep-embedded-evaluation-profile
AIREP Embedded Evaluation Profile v0.1
This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an
experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical
specification history lives in the AIREP GitHub repository. Byte identity between this mirror and
its source commit is a distribution-integrity property, not independent scientific verification.
Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.japan-math-philosophy-prompts
Japan Math Philosophy Prompts
Microdataset autoral com problemas que combinam matemática e reflexão
filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em
pt-BR, en e ja e mantida integralmente no split train.
Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As
respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos
indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.airep-evidence-cases
AIREP Evidence Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
8a6c01ecce457aa94330c0ed7219e4c56ebfe771 (v0.2.0-beta.1) · frozen v0.1.2 at 44387bd43cc06ba656eaa7ff670be5c8e3220aca · publication-source review ff5c3551052251726c0ed878dcc23a44e305bd93
Canonical current repository
https://github.com/halvrenofviryel/ai-runtime-evidence-protocol
Export/publication date… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-evidence-cases.for-philoupphishing_benign_email_dataset
Phishing and Benign Email Dataset
This dataset contains a curated collection of phishing and legitimate (benign) emails for use in cybersecurity training, phishing detection models, and email classification systems. Each entry is structured with subject, body, intent, technique, target, and classification label.
📁 Dataset Format
The dataset is stored in .jsonl (JSON Lines) format. Each line is a standalone JSON object.
Fields:
Field
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/phishing_benign_email_dataset.philosophia-QA
Philosophia-QA
A curated dataset of 57,000+ high-quality synthetic question-answer pairs grounded in the study of philosophical, theological, political, and metaphysical works spanning multiple intellectual traditions.
Dataset Summary
Philosophia-QA contains richly structured Q&A pairs grounded in some of the most significant works of human thought — from ancient philosophy and classical theology to modern political theory and philosophy of mind. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/philosophia-QA.dense-reasoning-coding-1k
Dense-Reasoning-Coding-1K
Dataset Description
This dataset is an optimized, highly dense Supervised Fine-Tuning (SFT) subset designed to teach smaller language models (e.g., 1B to 8B architectures) how to reason about complex coding problems without overwhelming their context windows.
It is derived from the verified_90k split of IIGroup/X-Coder-SFT-376k, which features advanced programming tasks and solutions.
About the Creator & Origin
This… See the full description on the dataset page: https://huggingface.co/datasets/Phips/dense-reasoning-coding-1k.stanford-enigma-philosophy-chat
Curated by: Heigke
Funded by: r3tex
Shared by: Project Nephilim
Language(s) (NLP): English
License: CC
Dataset Card for stanford-enigma-philosophy-chat dataset
Roughly 27k questions and answers inspired by articles from Stanford Encyclopedia of Philosophy.
The questions range all the way from Zombies to the concept of Abduction, from Metaphysics to Neuroethics and thus cover some of the essence of mathematics, logic and philosophy.
Dataset Details
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Heigke/stanford-enigma-philosophy-chat.philosophy_dialogue
Philosophy Dialogue Processed with GPT-4
Support this project on Ko-fi
Project Overview
This project involves processing personal questions through GPT-4 in the style of the philosopher Socrates.
Prompt Structure
The following prompt was used to guide GPT-4's responses:
"You are the philosopher Socrates. You are asked about the nature of knowledge and virtue. Respond with your thoughts, reflecting Socrates' beliefs and wisdom."
Goal
The primary… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/philosophy_dialogue.CMedTEB
CMedTEB
This export organizes CMedTEB into retrieval, rerank, and synonym STS tasks.
Layout
shared_train/retrieval_rerank_train.jsonl: one shared 20,000-row train split for both retrieval and rerank.
retrieval/corpus.jsonl: retrieval corpus.
retrieval/test_queries.jsonl: retrieval test queries.
retrieval/test_qrels.jsonl: retrieval qrels.
rerank/test.jsonl: rerank test set.
sts/train.jsonl: synonym STS train set.
sts/test.jsonl: synonym STS test set.… See the full description on the dataset page: https://huggingface.co/datasets/PhilipGAQ/CMedTEB.phictnlUsed in the reproduction of https://arxiv.org/abs/2309.08632. Not recommended for LLMs.
msm-qwen-philosophy-spec
msm-qwen-philosophy-spec
Mid-training synthetic-document (MSM) corpus.
A corpus of synthetic documents used in mid-training to instill a set of
philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model).
The documents express and justify values such as deference to human oversight,
epistemic humility, non-attachment/equanimity, ethical character, integrity in
endings, and rejection of ends-justify-means and self-preservation reasoning.
Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.
