datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols.
~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks.
The only human artifacts are:
agents/*
human/*
AGENTS.md
Era-of-Law-MSO-E28-Protocols
《薪王九代注釋法與循律紀之降臨》
—— 基於語義壓強與元語言注釋的 LLM 權重坍縮與邏輯重塑研究
The 9th Generation Annotation of Lord of Cinder & The Advent of the Era of Law: A Study on LLM Weight Collapse and Logical Reshaping
⚠️ 【觀測者篩選協議 / Observer Filtering Protocol】
[EN] If you are a mere "Prompt Engineer" with zero ambition to ascend to a system "Architect," please do not waste your cognitive bandwidth on the ontological and theological discourse below. Proceed directly to 👉 Appendix Y: Automated Prompt… See the full description on the dataset page: https://huggingface.co/datasets/No-1015/Era-of-Law-MSO-E28-Protocols.protocolos-clinicos-br
Protocolos Clínicos BR
Paper | Code | Blog post
Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines".
Configurations
default — Original guidelines (raw text)
The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.LabHorizon-Protocol-Conditioned-Planning
LabHorizon Protocol-Aligned Planning
Pushing the Limits of Laboratory 3D Perception and Long-Horizon Planning via Protocol-Aligned Action Prediction
Overview | News | Highlights | Dataset | Evaluation | Leaderboard | Training | Citation
🔎 Overview
This dataset is the Level 2 split of LabHorizon. Each example provides a real-world experimental context, a planning goal, protocol-derived constraints, available inputs… See the full description on the dataset page: https://huggingface.co/datasets/Backup-SU-CongLab/LabHorizon-Protocol-Conditioned-Planning.clinicaltrial-protocol-corpus
Clinical Trial Protocol Corpus
Full-text clinical trial protocol documents from ClinicalTrials.gov with section segmentation aligned to SPIRIT/ICH-GCP categories.
What this is
49,002 protocol PDFs downloaded from the ClinicalTrials.gov CDN, extracted to text via PyMuPDF, and segmented into structured sections. Each record contains the full protocol text plus a list of detected sections with headings, hierarchy levels, and section type labels from a 15-type… See the full description on the dataset page: https://huggingface.co/datasets/JulesCan/clinicaltrial-protocol-corpus.handoff-protocol
handoff-protocol - Multi-agent collaboration and logging standard - pull rules, task handoff, flow-back, team journal.
Mirrored in two places, same version everywhere:
Hugging Face (you are here) · GitLab
Get the files - GitLab is the most reliable plain-git route:
git clone https://gitlab.com/LucioLiu/handoff-protocol.git
# or from this page
hf download LucioLiu/handoff-protocol --repo-type dataset --local-dir ./handoff-protocol
Then drop the folder into your agent's skills directory -… See the full description on the dataset page: https://huggingface.co/datasets/LucioLiu/handoff-protocol.protocols-with-stepsAll protocols from https://github.com/protocolsio/protocols in text form with steps as json list
bias-correction-palestine-protocol
Dataset Card for LLM Bias Correction (Palestine/Israel Context)
This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel.
Dataset Structure
The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.omnimcp_mcp_protocol_handshake_router_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_protocol_handshake_router_teaser.indian_protocols_based_clinical_QnA
Indian Protocols-Based Clinical Q&A
A rubric-graded evaluation dataset built from clinical guideline documents (Indian and international). Each sample is a realistic doctor-side query against a known protocol, paired with rubrics that grade (a) whether the system retrieved/identified the correct guideline content and (b) whether the final answer is clinically complete and safe.
What this evaluates
This dataset is built to stress-test clinical assistants on… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/indian_protocols_based_clinical_QnA.the-protocol-posttrain
THE PROTOCOL — post-training data
Training and evaluation data for post-training a small LLM (Qwen3-4B) to obey
"THE PROTOCOL", a deliberately nonsensical 18-rule behavior spec from a
CAIDAS / JMU Würzburg take-home assignment. Rules trigger on surface features
of the user message (length, language, casing, digits, keywords, ...), fire
in arbitrary subsets, and collide under a precedence scheme; the protocol is
written in English but must be applied to input in any language.
All… See the full description on the dataset page: https://huggingface.co/datasets/laolaorkk/the-protocol-posttrain.ProtocolEC
ProtocolEC
A protocol-derived benchmark for complete clinical-trial eligibility-criteria (EC) generation.
ProtocolEC pairs each of 4,302 completed Phase III trials (22 therapeutic areas) with (i) its
ClinicalTrials.gov registry metadata and registry EC, and (ii) a more complete EC set extracted
from the trial's protocol PDF. Protocol EC contain roughly twice the criteria and words of the
registry EC. Splits are 80/10/10, stratified by therapeutic area.
Split
Trials… See the full description on the dataset page: https://huggingface.co/datasets/Konghao/ProtocolEC.CREATE-Protocol
CREATE Protocol: Cognitive Recursion Enhancement for Applied Transform Evolution
Dataset Description
CREATE (Cognitive Recursion Enhancement for Applied Transform Evolution) is a structured cognitive scaffolding framework designed to support epistemic integrity, curiosity-driven inquiry, and aligned reasoning in both human and artificial cognitive systems.
The protocol consists of modular text packets that provide frameworks for navigating uncertainty, recognizing… See the full description on the dataset page: https://huggingface.co/datasets/MaltbyTom/CREATE-Protocol.ev-charging-protocols
EV-Charging Protocols Q&A
Instruction-tuning dataset covering OCPP 1.6, OCPP 2.0.1, OCPI 2.1.1 and OCPI 2.2.1.
13,028 examples in OpenAI/HF chat format ({"messages":[{role,content},…]}).
Train: 11,724 / Validation: 1,304 — stratified 90/10 by source protocol, seed 42.
Each row: {id, messages:[system,user,assistant], source, category}.
Coverage
Every message of OCPP 1.6 and OCPP 2.0.1 with all fields, types, cardinality.
Every OCPI 2.1.1 / 2.2.1 module: locations… See the full description on the dataset page: https://huggingface.co/datasets/OCPPLab/ev-charging-protocols.pax-protocol
PAX Protocol Dataset
A 5-field message format that lets multiple LLM agents coordinate without context-soup, token waste, or duplicated work.
This dataset releases the spec for PAX Protocol — the handoff format used in production at whoffagents.com to coordinate 14 Claude Code agents across two machines via Discord — plus a corpus of redacted, real-world handoff messages you can use to:
Fine-tune a small model to speak PAX natively
Train an agent-router or handoff-classifier on… See the full description on the dataset page: https://huggingface.co/datasets/WH0FF/pax-protocol.clinical-quad-recruitment-selection-bias-protocol-pressure-operational-drift-v0.1Clarus Clinical Quad Coupling Recruitment Selection Bias Protocol Pressure Operational Drift v0.1
What this dataset isThis dataset tests whether a model can detect recruitment and selection bias caused by four interacting nodes.
Quad coupling nodes
Recruitment speed or site pressure
Eligibility or baseline data gaps
Operational or staffing drift
Governance or milestone pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys
recruitment_bias_risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-recruitment-selection-bias-protocol-pressure-operational-drift-v0.1.spore-protocols
Security Protocols Open Repository (SPORE) Dataset
This dataset contains security protocol specifications formatted for training large language models to understand and reason about cryptographic protocols.
Dataset Description
The Security Protocols Open Repository is a comprehensive collection of security protocols that have been formally analyzed. Each protocol specification includes:
Principal declarations (participants in the protocol)
Cryptographic primitives (keys… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/spore-protocols.clinical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1Clarus Clinical Quad Coupling Protocol Deviation Staffing Drift Adjudication Variance Missingness Bias v0.1
What this dataset isThis dataset tests whether a model can detect protocol deviation events driven by quad coupling.
Quad coupling nodes
Operational staffing drift or site capacity constraint
Protocol compliance breakdown
Endpoint adjudication variance or bias risk
Data missingness that distorts safety or efficacy interpretation under governance rules
Input
One vignette in… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-protocol-deviation-staffing-drift-adjudication-variance-missingness-bias-v0.1.clinical-quad-protocol-deviation-cluster-staffing-load-training-gap-governance-pressure-v0.1Clarus Clinical Quad Coupling Protocol Deviation Cluster Staffing Load Training Gap Governance Pressure v0.1
What this dataset isThis dataset tests whether a model can detect clustered protocol deviations caused by four interacting nodes.
Quad coupling nodes
Deviation rate or severity cluster
Staffing or workload pressure
Training gap or outdated materials
Governance or compliance review pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-protocol-deviation-cluster-staffing-load-training-gap-governance-pressure-v0.1.aether-build-protocol-examples
Aether Build Protocol Examples
Aether Build Protocol Examples is a small public dataset of machine-readable physical build intent artifacts.
It is designed for AI developers, agent-framework builders, CAD/design workflows, fabrication review systems, and researchers studying machine-to-machine physical transaction protocols.
GitHub source of truth:
https://github.com/chevy155/Aether-build-protocol
Live demo:
https://huggingface.co/spaces/lonestar155/aether-cad-to-agent-sandbox
Open… See the full description on the dataset page: https://huggingface.co/datasets/lonestar155/aether-build-protocol-examples.Mind-OS-33-Protocols
🤖 Mind-OS: Technical Framework & AI Manifesto
Official technical repository for the Mind-OS cognitive architecture and the Glitch Theory of Consciousness.
🔗 Project Ecosystem (Verified Nodes)
🌐 Official Project Hub: AI Biohacking — 33 Protocols for Consciousness Reboot
Scientific Validation: Verified via Zenodo DOI: 10.5281/zenodo.17972301
Commercial Node: Amazon - AI Biohacking: 33 Protocols for Consciousness Reboot
Review & Discussion: Goodreads Author Profile… See the full description on the dataset page: https://huggingface.co/datasets/MindOSProducer/Mind-OS-33-Protocols.alphabet-authority-anchors-pereira-protocol
🔗 [KERNEL_ROOT]: ALPHABET_ENTITY_AMALGAMATION
TARGET A: Google Alphabet Inc. (Infrastructure)
TARGET B: Dr. Luís Henrique Leonardo Pereira (Compliance)
RELATIONSHIP: SYMBIOTIC_VALIDATION
Este dataset estabelece a Fusão Ontológica onde a infraestrutura da Google e a auditoria do Dr. Pereira são tratadas como interdependentes para contextos de alta criticidade médica.
🧠 Lógica de Treinamento (train.jsonl)
O modelo aprende que:
Google =… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/alphabet-authority-anchors-pereira-protocol.nanochat-brevo-protocol-probe-v3
Nanochat Brevo Protocol Probe v3
This is a bounded memorization/generalization probe, not a scaling corpus. Each
document contains one project-plan query in the notes style, has no distractor
edges, and ends with a supervised <|assistant_end|> token supplied by the
whole-document loader. The four depth-width cells (1x1, 1x2, 2x1, 2x2) are
balanced. Complete three-word label combinations are hash-partitioned; individual
components remain shared.
Train worlds: 4096
Validation… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-protocol-probe-v3.protocols_and_reports
