datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.hipaa-compliance-training
HIPAA Compliance Training Dataset
Dataset Description
The first comprehensive HIPAA compliance training dataset for LLM fine-tuning, covering the Security Rule, Privacy Rule, Breach Notification Rule, and implementation guidance from NIST and FDA.
Dataset Summary
Total Examples: 1,287 (1,029 train / 258 validation)
Source Documents: 9 federal publications (~5.6 MB extracted content)
Format: JSONL with chat-formatted messages
License: CC0-1.0 (Public Domain)… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/hipaa-compliance-training.iclr2026-lm-logprobs
LM Log-Probabilities for Value Bias Analysis
Next-token log-probability distributions from 12 language models across 54 prompts, used in the paper:
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A.F. Thompson, Elle, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, Tsvetomira Dumbalska (ICLR 2026)
Part of the Oxford-HIPlab collection for this paper.
Dataset description
Each CSV contains the full next-token log-probability… See the full description on the dataset page: https://huggingface.co/datasets/Oxford-HIPlab/iclr2026-lm-logprobs.omnimcp_healthtech_hipaa_redactor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_hipaa_redactor_teaser.hipaa_dental_tickets
Hipaa_Dental_Tickets (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating HIPAA Dental Office Support Tickets for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and validation… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/hipaa_dental_tickets.HippoTarget
🎯 HippoTarget
A curated drug-target interaction dataset, teaching LLMs which molecules bind to which proteins.
💡 Overview
Welcome to HippoTarget, the fifth member of the ZemResearch Hippo Ecosystem. Before a drug can do anything useful in the body, it first has to bind to the right protein — like a key fitting into a lock. HippoTarget teaches LLMs exactly that: given a small molecule, which protein does it interact with?
This dataset combines real… See the full description on the dataset page: https://huggingface.co/datasets/ZemResearch/HippoTarget.HippoSynth
⚗️ HippoSynth
A curated, reaction-ready dataset for teaching LLMs the art of chemical synthesis.
💡 Overview
Welcome to HippoSynth, the third member of the ZemResearch Hippo Ecosystem. While HippoCrates teaches LLMs what molecules look like, HippoSynth teaches them how molecules are made.
This dataset covers the full spectrum of chemical synthesis tasks — from predicting reaction products given a set of reactants (forward synthesis), to working backwards… See the full description on the dataset page: https://huggingface.co/datasets/ZemResearch/HippoSynth.hospital-twin-validation-200
HipAAsynth Dataset
Summary
This dataset is a validation artifact generated by HipAAsynth.
HipAAsynth is a deterministic testing and validation service that simulates real-world variability to evaluate how healthcare systems perform under deployment conditions.
Description
This dataset represents a controlled cohort used for testing and benchmarking.
HipAAsynth generates cohorts to simulate how conditions present across:
patient populations
demographic… See the full description on the dataset page: https://huggingface.co/datasets/HipAAsynth/hospital-twin-validation-200.chest-pain-diagnostic-cohort-100
HipAAsynth Dataset
Summary
This dataset is a validation artifact generated by HipAAsynth.
HipAAsynth is a deterministic testing and validation service that simulates real-world variability to evaluate how healthcare systems perform under deployment conditions.
Description
This dataset represents a controlled cohort used for testing and benchmarking.
HipAAsynth generates cohorts to simulate how conditions present across:
patient populations
demographic… See the full description on the dataset page: https://huggingface.co/datasets/HipAAsynth/chest-pain-diagnostic-cohort-100.HippoXic
☠️ HippoXic
A premium instruction-tuning dataset for teaching LLMs chemical toxicology, clinical safety, and side effects.
💡 Overview
Welcome to HippoXic, curated by ZemResearch. If our previous dataset (HippoCrates) taught AI how to generate molecules, this dataset teaches AI how to keep us safe from them.
We combined three legendary bio-informatics databases into one streamlined, chat-ready format:
Tox21: For biological toxicity pathways (Does it… See the full description on the dataset page: https://huggingface.co/datasets/ZemResearch/HippoXic.HippoLv
💊 Hippolv
A highly curated, instruction-tuning dataset for teaching AI how drugs behave in the human body (ADMET & Solubility).
💡 Overview
Welcome to Hippolv, the third specialized dataset curated by ZemResearch. While our previous datasets focused on general molecular structures and toxicology, Hippolv is engineered to bridge the gap between pure chemistry and clinical pharmacology.
This dataset focuses on ADMET (Absorption, Distribution, Metabolism… See the full description on the dataset page: https://huggingface.co/datasets/ZemResearch/HippoLv.HippoCrates
🧬 HippoCrates
A hyper-sterilized, ready-to-use molecular dataset for fine-tuning Chemical LLMs.
💡 Overview
Welcome to HippoCrates, curated by ZemResearch. This dataset is specifically designed to teach Large Language Models (LLMs) the universal language of chemistry: SMILES.
We extracted raw molecular data from the legendary QM9 (for basic organic structures) and ChEMBL (for bioactivity and medicinal characteristics) databases, and transformed them… See the full description on the dataset page: https://huggingface.co/datasets/ZemResearch/HippoCrates.mejores-hipotecas-2025Fuente: https://www.finect.com/articulos/mejores-hipotecasAutor: FinectLicencia: CC-BY-NC 4.0
