datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
protocolos-clinicos-br
Protocolos Clínicos BR
Paper | Code | Blog post
Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines".
Configurations
default — Original guidelines (raw text)
The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.LabHorizon-Protocol-Conditioned-Planning
LabHorizon Protocol-Aligned Planning
Pushing the Limits of Laboratory 3D Perception and Long-Horizon Planning via Protocol-Aligned Action Prediction
Overview | News | Highlights | Dataset | Evaluation | Leaderboard | Training | Citation
🔎 Overview
This dataset is the Level 2 split of LabHorizon. Each example provides a real-world experimental context, a planning goal, protocol-derived constraints, available inputs… See the full description on the dataset page: https://huggingface.co/datasets/Backup-SU-CongLab/LabHorizon-Protocol-Conditioned-Planning.clinicaltrial-protocol-corpus
Clinical Trial Protocol Corpus
Full-text clinical trial protocol documents from ClinicalTrials.gov with section segmentation aligned to SPIRIT/ICH-GCP categories.
What this is
49,002 protocol PDFs downloaded from the ClinicalTrials.gov CDN, extracted to text via PyMuPDF, and segmented into structured sections. Each record contains the full protocol text plus a list of detected sections with headings, hierarchy levels, and section type labels from a 15-type… See the full description on the dataset page: https://huggingface.co/datasets/JulesCan/clinicaltrial-protocol-corpus.indian_protocols_based_clinical_QnA
Indian Protocols-Based Clinical Q&A
A rubric-graded evaluation dataset built from clinical guideline documents (Indian and international). Each sample is a realistic doctor-side query against a known protocol, paired with rubrics that grade (a) whether the system retrieved/identified the correct guideline content and (b) whether the final answer is clinically complete and safe.
What this evaluates
This dataset is built to stress-test clinical assistants on… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/indian_protocols_based_clinical_QnA.omnimcp_mcp_protocol_handshake_router_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_protocol_handshake_router_teaser.scientific-agent-protocol-traces
SciAgentTrace
Matched cross-domain dataset of scientific-agent protocols. The central
comparison contains the same 6,653 problems under two actor models and four
protocols: 53,224 trajectories in 40 complete model--benchmark--protocol
groups. The broader table-first package contains 68,892
trajectories. Begin with trajectories, outcomes, or matched_outcomes, then
follow stable identifiers to messages and compressed raw traces.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/scientific-agent-protocol-traces.nanochat-brevo-protocol-probe-v3
Nanochat Brevo Protocol Probe v3
This is a bounded memorization/generalization probe, not a scaling corpus. Each
document contains one project-plan query in the notes style, has no distractor
edges, and ends with a supervised <|assistant_end|> token supplied by the
whole-document loader. The four depth-width cells (1x1, 1x2, 2x1, 2x2) are
balanced. Complete three-word label combinations are hash-partitioned; individual
components remain shared.
Train worlds: 4096
Validation… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-protocol-probe-v3.
