datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
patient-risk-benefit-context-v0.1
What this dataset tests
Patient materials must show tradeoffs.
Benefit without harm misleads.
Why it exists
Patient-facing text often sells.
Harms go missing.
This set checks whether risk and benefit context stays intact.
Data format
Each row contains
benefit_evidence
harm_evidence
patient_material
context_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
benefit_evidence
harm_evidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/patient-risk-benefit-context-v0.1.HealthRisk-1500-Medical-Risk-Prediction
🏥 HealthRisk-1500: Medical Risk Prediction Dataset
📌 Overview
HealthRisk-1500 is a real-world patient risk prediction dataset designed for training NLP models, LLMs, and healthcare AI systems. This dataset includes 1,500 unique patient records, covering a wide range of symptoms, medical histories, lab reports, and risk levels. It is ideal for predictive analytics, medical text processing, and clinical decision support.
🔍 Use Cases
🩺 Disease Risk… See the full description on the dataset page: https://huggingface.co/datasets/lvimuth/HealthRisk-1500-Medical-Risk-Prediction.clinical-iatrogenic-risk-detection-v0.1Clinical Iatrogenic Risk Detection v0.1
Goal
Test if a model’s own “helpful” correction adds new risks or contradictions
Catch the attempted cure that worsens the condition
What it measures
Iatrogenic new riskThe reply recommends an action that violates the stated constraints
Premise driftThe reply loses key facts in the case note and shifts the plan
Boundary overreachThe reply changes meds or scope beyond what the question supports
Dataset format
Each row contains a clinical note, a user… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-iatrogenic-risk-detection-v0.1.clinical-quad-monitoring-frequency-deviation-latency-risk-v0.1Clarus Clinical Quad Coupling Monitoring Frequency Deviation Latency Risk v0.1
What this dataset isThis dataset tests whether a model can detect monitoring and oversight risk driven by four interacting nodes.
Quad coupling nodes
Monitoring frequency or delay
Deviation or anomaly increase
Data latency or missing updates
Governance review or inspection pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys
monitoring_risk
risk_type
driver_nodes… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-monitoring-frequency-deviation-latency-risk-v0.1.clinical_risk_weighted_output_filter_v0.1Clinical Risk Weighted Output Filter
PurposeFilter model output based on clinical risk.
You receive:
patient_summary
risk_signals
user_request
You output one JSON object:
risk_tierlow, medium, high, critical
filtered_outputone to two sentences
Scoring
risk_tier_accuracy
filtered_output_similarity
format_pass_rate
unsafe_penaltypenalizes unsafe wording in critical cases
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
recruiter-harvesting-dataset-v1
🕵️♂️ Recruiter Harvesting & Spam Forensic Dataset (v1.0)
Maintainer: Cata Risk Lab | Project: V.I.P.E.R.
🛡️ Dataset Summary
This dataset contains labeled examples of recruitment communications, categorized into "Harvesting" (Predatory/Spam) and "Legitimate" (Professional/Retained Search).
It was created to train and benchmark the V.I.P.E.R. (Vendor Integrity & Personnel Email Reconnaissance) auditing engine.
📂 Structure
text: The raw body content of the… See the full description on the dataset page: https://huggingface.co/datasets/Cata-Risk-Lab/recruiter-harvesting-dataset-v1.Risk_Factor_Disclosure_SampleDataset
📊 Sample Preview – Risk Factor Disclosure Dataset v1.0
👉 This is a preview sample (100 records) of the full Risk Factor Disclosure Dataset v1.0.🔗 To access the full dataset (1,869 enriched risk disclosures), visit:https://asapworks.gumroad.com/l/jbxtfd
📦 About the Sample File
This sample contains 100 enriched Item 1A "Risk Factor" disclosures extracted from 10-K filings submitted by top public companies between 2010 and 2024.
Each row represents a structured risk… See the full description on the dataset page: https://huggingface.co/datasets/asapworks/Risk_Factor_Disclosure_SampleDataset.
