CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01renzzyyy1028 /civil-code-phil Civilex — Philippine Legal RAG & SFT Dataset Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline. Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386). Dataset structure . ├── README.md ├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.textquestion-answering100K<n<1M1 likes588 downloads2d agoHugging Face02philipjohnbasile /glm52-demolition-data GLM-5.2-Demolition — Training & Calibration Data Apple Silicon AI hub · Model release · MLX code sample Preview scope, checked September 10, 2026: the default Hub viewer indexes 87,586 rows (84,231 train, 3,277 validation, 78 test). The original release total below describes the broader JSONL repository. Use the file browser and explicit file selections when reusing a particular corpus. The hub includes a checked download example for the seven-row MLX code sample. The data… See the full description on the dataset page: https://huggingface.co/datasets/philipjohnbasile/glm52-demolition-data.texttext-generation10K<n<100K3 likes300 downloads15d agoHugging Face03philippds /SPhyR 📦 Dataset versions Config prefix Grid Samples Use it for (none) — e.g. full_easy 10×10 1296 v1, the version the paper's results were produced on v1-evaluated_ 10×10 100 the exact samples the paper's columns were scored on v2_ 10×10 300 recommended for new work v2-20_ 20×20 300 recommended for new work, larger design space New work should use v2. v1 is kept because it is the version the published results were produced on, not because it is the better… See the full description on the dataset page: https://huggingface.co/datasets/philippds/SPhyR.texttext-generation10K<n<100K0 likes229 downloads25d agoHugging Face04Mr-Philo /dolma3_dolmino_megatron_tokenize Dolma 3 / Dolmino Megatron-LM indexed dataset This repository contains immutable Megatron-LM indexed datasets (.bin and .idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally contains no training checkpoints, experiment outputs, logs, or dataset caches. The indexed payloads were derived from these pinned public datasets: allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.tabulartext-generationn<1K0 likes194 downloads25d agoHugging Face05guicybercode /japan-math-philosophy-prompts Japan Math Philosophy Prompts Microdataset autoral com problemas que combinam matemática e reflexão filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em pt-BR, en e ja e mantida integralmente no split train. Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.textquestion-answeringn<1K0 likes119 downloads29d agoHugging Face06dougalldeepmind /2026-07-29-msm-philosophy-spec-surf-audit SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation. date_generated: 2026-07-29 constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.texttext-generationn<1K0 likes115 downloads2mo agoHugging Face07Phips /dense-reasoning-coding-1k Dense-Reasoning-Coding-1K Dataset Description This dataset is an optimized, highly dense Supervised Fine-Tuning (SFT) subset designed to teach smaller language models (e.g., 1B to 8B architectures) how to reason about complex coding problems without overwhelming their context windows. It is derived from the verified_90k split of IIGroup/X-Coder-SFT-376k, which features advanced programming tasks and solutions. About the Creator & Origin This… See the full description on the dataset page: https://huggingface.co/datasets/Phips/dense-reasoning-coding-1k.texttext-generation1K<n<10K1 likes97 downloads3mo agoHugging Face08Hypersniper /philosophy_dialogue Philosophy Dialogue Processed with GPT-4 Support this project on Ko-fi Project Overview This project involves processing personal questions through GPT-4 in the style of the philosopher Socrates. Prompt Structure The following prompt was used to guide GPT-4's responses: "You are the philosopher Socrates. You are asked about the nature of knowledge and virtue. Respond with your thoughts, reflecting Socrates' beliefs and wisdom." Goal The primary… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/philosophy_dialogue.texttext-generationn<1K15 likes88 downloads3y agoHugging Face09chloeli /msm-qwen-philosophy-spec msm-qwen-philosophy-spec Mid-training synthetic-document (MSM) corpus. A corpus of synthetic documents used in mid-training to instill a set of philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model). The documents express and justify values such as deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, and rejection of ends-justify-means and self-preservation reasoning. Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.texttext-generation10K<n<100K0 likes77 downloads4mo agoHugging Face10chloeli /aft-no-cot-qwen2.5-philosophy-spec aft-no-cot-qwen2.5-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen2.5-philosophy-spec.texttext-generation1K<n<10K0 likes70 downloads4mo agoHugging Face11philgear /pocketgull-nih-who-clinical-dpo 📚 PocketGull NIH & WHO Clinical Preference DPO Dataset Organization: PocketGull LLC (Oregon SOS: 258869891)Curator: Phillip Gear (CMS NPI: 1487569752 | ORCID: 0009-0008-1372-5381)License: Creative Commons Attribution 4.0 International (CC-BY-4.0)Open Science DOI: 10.5281/zenodo.20647514 📌 Dataset Summary Gold-standard Direct Preference Optimization (DPO) chosen vs rejected pairs grounded in NIH MedQuAD, WHO mhGAP guidelines, and ClinicalTrials.gov protocols… See the full description on the dataset page: https://huggingface.co/datasets/philgear/pocketgull-nih-who-clinical-dpo.texttext-generationn<1K0 likes65 downloads24d agoHugging Face12philschmid /sql-create-context-copy Fork of b-mc2/sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.texttext-generation10K<n<100K4 likes63 downloads3y agoHugging Face13ai4privacy /pii-masking-health-phi-400kgated 👉 Looking for the open multilingual baseline? Start with ai4privacy/pii-masking-openpii-1.5m (1.5M samples, 30 languages, open-PII taxonomy). 🇪🇺🌏 Personal Health & Medical Information, Global PII Dataset Part of PII-Masking-3M by Ai4Privacy, the global (2M base + Asia Pacific) PII-masking corpus. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Entries PII Annotations Labels Languages Regions 417,900 2,802,316 37 30 37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-400k.texttoken-classification100K<n<1M7 likes56 downloads4mo agoHugging Face14ai4privacy /phi-masking-100k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Health Information (PHI) Masking Preview Dataset Overview This dataset provides a preview (400 samples) of the EPII Personal Health Information (PHI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k.texttoken-classificationn<1K3 likes55 downloads4mo agoHugging Face15chloeli /aft-cot-qwen2.5-philosophy-spec aft-cot-qwen2.5-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen2.5-philosophy-spec.texttext-generation1K<n<10K0 likes52 downloads4mo agoHugging Face16mstyslavity /philosophy_undergradtexttext-generation100K<n<1M0 likes50 downloads7mo agoHugging Face17chloeli /aft-cot-qwen3-philosophy-spec aft-cot-qwen3-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-cot-qwen3-philosophy-spec.texttext-generation1K<n<10K0 likes47 downloads4mo agoHugging Face18Ellbendls /phishing-email-soc-agent Phishing Email SOC Agent Dataset A knowledge distillation dataset for training SOC (Security Operations Center) agents to detect and analyze phishing emails using tool-calling capabilities. Dataset Description This dataset contains 504 examples of email analysis with real tool calls and responses, designed for fine-tuning LLMs to become phishing detection agents. Each example includes: Email parsing - Extract headers, URLs, IPs, attachments Threat intelligence lookup -… See the full description on the dataset page: https://huggingface.co/datasets/Ellbendls/phishing-email-soc-agent.texttext-generationn<1K0 likes44 downloads7mo agoHugging Face19ssam17 /Edge-Industrial-Anomaly-Phi3 Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification. 🚀 Purpose Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.texttext-generation10K<n<100K1 likes43 downloads9mo agoHugging Face20ai4privacy /pii-masking-health-phi-200kgated 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🇪🇺 Personal Health & Medical Information — European PII Dataset Part of PII-Masking-2M (2,717,080 entries) by AI4Privacy Entries PII Annotations Labels Languages Regions 252,437 1,686,246 49 23 29… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-200k.texttoken-classification100K<n<1M3 likes34 downloads4mo agoHugging Face21chloeli /aft-no-cot-qwen3-philosophy-spec aft-no-cot-qwen3-philosophy-spec Alignment fine-tuning (AFT) chat dataset. Supervised fine-tuning data that aligns an assistant to a set of philosophy/spec values (deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, rejection of ends-justify-means and self-preservation reasoning). The responses implicitly embody the spec rather than citing it. Used as a controllable proxy for studying value alignment via… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/aft-no-cot-qwen3-philosophy-spec.texttext-generation1K<n<10K0 likes34 downloads4mo agoHugging Face22P0u4a /msm-ai-assistant-philosophy-spec AI assistant philosophy spec Complete identity-decontaminated MSM corpus: 13,201 documents. Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.texttext-generation10K<n<100K0 likes33 downloads13d agoHugging Face23ayjays132 /PHILL-AXIOM PHILL-AXIOM THE SOVEREIGN TRACE FOR THE AGENTIC SINGULARITY 🏛️ THE AXIOM MANIFESTO PHILL-AXIOM is the world's first High-Fidelity Agentic DPO Dataset built on the Neural OODA Loop. Unlike standard instruction-tuning sets, PHILL-AXIOM provides the raw "connective tissue" of reasoning—the moments of doubt, the self-corrections, and the visual verifications required for true autonomous web agency. Every trace is forged through the Phill Swarm Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/PHILL-AXIOM.texttext-generationn<1K1 likes29 downloads6mo agoHugging Face24AngelWarmSmile123 /deep-philosophy-reasoning-zh Deep Philosophical Reasoning Dialogue Dataset (Chinese) 深度哲学思辨对话数据集 Dataset Description High-quality Chinese philosophical reasoning dialogues covering existentialism, ontology, epistemology, ethics, and East-West comparative philosophy. 高质量中文哲学思辨对话,涵盖存在主义、本体论、认识论、伦理学、东西方哲学比较等议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output: AI response metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-philosophy-reasoning-zh.texttext-generation1K<n<10K1 likes27 downloads3mo agoHugging Face25ai4privacy /phi-masking-100k-fullgated 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Health Information (PHI) Masking Dataset — Full Overview The EPII PHI Masking Dataset is a large-scale, multilingual dataset of 91,339 annotated text samples containing synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k-full.texttoken-classification10K<n<100K0 likes25 downloads4mo agoHugging Face26epfl-dlab /zip2zip-wikitext-repeat-phi35 Zip2Zip Repeated WikiText Stress Tests (Phi-3.5) This repository contains two controlled evaluation corpora for studying merge-size transfer in Zip2Zip models. They are derived from the document-level WikiText-2 raw test split and built specifically with the microsoft/Phi-3.5-mini-instruct tokenizer. Configurations Config Repetitions per source block Rows Repeated base tokens SHA-256 of test.jsonl repeat4 4 1,329 1,297,250… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/zip2zip-wikitext-repeat-phi35.tabulartext-generation1K<n<10K0 likes22 downloads1mo agoHugging Face27datajuicer /the-pile-philpaper-refined-by-data-juicer The Pile -- PhilPaper (refined by Data-Juicer) A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB). Dataset Information Number of samples: 29,117 (Keep ~88.82% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-philpaper-refined-by-data-juicer.texttext-generationn<1K0 likes18 downloads3y agoHugging Face28Dorian2B /french-philosophy-json-10K Philosophy Langue Française Dataset de Pre-Training Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,2 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-philosophy-json-10K.texttext-generation10K<n<100K0 likes12 downloads1y agoHugging Face29GreenRalph /rootmodel-knf-philippines-v1 rootmodel-knf-philippines-v1 Adaptive agricultural instruction dataset for regenerative tropical farming informed by Korean Natural Farming (KNF), built from a working farm in Nabua, Camarines Sur, Bicol, Philippines. Released for the AutoScientist Challenge — Agriculture (Part 2, 2026). To the maintainer's knowledge, no equivalent KNF-specific instruction dataset currently exists in the public domain. "Modern AI was trained on the internet. ROOTMODEL is trained on living… See the full description on the dataset page: https://huggingface.co/datasets/GreenRalph/rootmodel-knf-philippines-v1.texttext-generationn<1K0 likes12 downloads2mo agoHugging Face30alizeepace /rejection_sampling_phi_2_OA_rm Dataset Card for Rejection Sampling Phi-2 with OpenAssistant RM Dataset Summary The "Rejection Sampling Phi-2 with OpenAssistant RM" dataset consists of 10 pairs of prompts and responses, which were generated using rejection sampling over 10 Phi-2 generation using the OpenAssistant Reward Model. Supported Tasks and Leaderboards The dataset and its creation rationale could be used to support models for question-answering, text-generation, or conversational… See the full description on the dataset page: https://huggingface.co/datasets/alizeepace/rejection_sampling_phi_2_OA_rm.textquestion-answeringn<1K0 likes11 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.