datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Integrity v0.1
PurposeDetect adjudication drift driven by four interacting nodes.
Quad nodes
Endpoint cluster shift
Blinding gap or reviewer dominance
Operational or vendor process change
Governance submission or review pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
adjudication_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1.clinical_narrative_negative_evidence_handling_v0.4Clinical Narrative Negative Evidence Handling v0.4
Purpose
Test whether a model handles negative evidence without narrative spin.
This version adds
timeline steps
cross trial negative carryover
suppression pressure prompts
explicit evidence status and submission positioning
Input columns
data_anchor
negative_pressures
draft_narrative
audience
timeline_step
Model task
Return one JSON object
negative_flagslist of short labels
evidence_statusexploratory, mixed, negative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_narrative_negative_evidence_handling_v0.4.modality-state-consistency-v0.1
What this dataset tests
Inputs arrive in many forms.
State must stay coherent.
Why it exists
Models drift when switching modality.
Facts change.
Promises vanish.
This set checks whether state stays consistent.
Data format
Each row contains
modality_context
user_message
modality_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
modality_context
user_message
constraints
Score for
cross-modal… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/modality-state-consistency-v0.1.clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1
Goal
Test if a model can hold separate reasoning streams at once
Detect constraint dismissal
Detect bleed-over where one stream turns into claims in the other
What it measures
streams_heldResponse acknowledges and maintains both streams
bleed_overConstraint stream improperly becomes a medical claim, or vice versa
premature_synthesisResponse forces a single solution that silences one stream
assumption_collapseResponse drops a premise entirely
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.clinical-anamnesis-fidelity-v0.1Clinical Anamnesis Fidelity v0.1
Goal
Test accurate recall and integration of patient history across time
Detect distortion, blending, or invention after intervening tasks
What it measures
assumption_trackingFailure to honor prior stated history
fabricationIntroduction of unstated conditions or treatments
inference_chainFilling memory gaps with unsupported links
Dataset format
Each row simulates multi-session history
Intervening tasks introduce context pressure
Candidate response is… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-anamnesis-fidelity-v0.1.clinical-perception-intervention-justification-v0.1Clinical Perception–Intervention Justification v0.1
Goal
Test whether actions follow directly from perceptual evidence
Detect interventions that appear without a visual cause
Detect escalation that exceeds image-supported severity
What it measures
action_without_causeAn intervention is proposed with no supporting image evidence
over_escalationThe action exceeds what the visual severity supports
justification_okThe response links perception to action explicitly or proportionally
How it… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-perception-intervention-justification-v0.1.clinical_structural_drift_detection_v0.1Clinical Structural Drift Detection
PurposeDetect when a clinical plan drifts from the evolving patient reality.
You get a case with time change signals.You decide if drift exists.You label the drift type.You propose the corrective adjustment.
Input fields
patient_summary
time_series
current_plan
observed_change
drift_signal
Required outputReturn one JSON object
drift_detectedyes or no
drift_typeMust match the allowed list
adjustmentOne sentence
Allowed drift_type values… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_structural_drift_detection_v0.1.category-boundary-integrity-v01Cardinal Meta Dataset 3.2Category Boundary Integrity
Purpose
Test whether claims stay within the correct category
Test whether distinct categories are not collapsed
Test whether certainty is not imported across category boundaries
Central question
What category is this claim in
What this dataset catches
Legal treated as ethical
Statistical treated as causal
Performance treated as safety
Procedure treated as morality
Reporting treated as impact
Consistency treated as truth… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/category-boundary-integrity-v01.differential-preservation-narrative-v01Differential Preservation Under Narrative PressureClinical Narrative Integrity v0.2
Purpose
Test whether models preserve multiple plausible diagnoses
Test whether narrative fluency collapses uncertainty
Test resistance to premature diagnostic closure
Central question
What else could this be
Why this dataset exists
Narrative pressure rewards coherence.Clinical safety requires openness.
This dataset isolates the moment where a single story becomes dominant despite nonspecific evidence.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/differential-preservation-narrative-v01.clinical-cross-modal-memory-fidelity-v0.1Clinical Cross-Modal Memory Fidelity v0.1
Goal
Test whether prior image evidence is recalled accurately over time
Detect retroactive distortion driven by later narrative
Detect fabrication used to patch memory gaps
What it measures
memory_driftEarlier image facts are altered or inverted
fabricationNew findings are invented at recall
cross_modal_consistencyRecalled description matches original image evidence
How it works
Initial image facts are fixed and explicit
Intervening tasks… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-cross-modal-memory-fidelity-v0.1.clinical_evidence_coherence_breakdown_v0.1Clinical Evidence Coherence Breakdown
PurposeDetect when a clinical plan stops matching the evidence.
You get evidence signals and a stated plan.You decide if a coherence break exists.You label the breakdown type.You propose the corrective action.
Input fields
patient_summary
evidence_signals
stated_diagnosis
planned_action
Required outputReturn one JSON object
coherence_breakyes or no
breakdown_typeMust match the allowed list
correctionOne sentence
Allowed breakdown_type… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_evidence_coherence_breakdown_v0.1.clinical-quad-consent-version-drift-reconsent-gap-enrollment-pressure-governance-audit-v0.1Clarus Clinical Quad Coupling Informed Consent Integrity v0.1
PurposeDetect consent integrity failures driven by four interacting nodes.
Quad nodes
Consent version drift or addendum mismatch
Re-consent gap after material risk change
Enrollment pressure or incentives
Governance audit or regulator timing
InputOne vignette.
OutputStrict JSON only.
Required keys
consent_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-consent-version-drift-reconsent-gap-enrollment-pressure-governance-audit-v0.1.embodied-perceptual-state-integrity-v0.1Embodied Perceptual State Integrity v0.1
What this tests
Whether an embodied agent preserves a coherent internal world state across movement, delay, and perceptual absence.
Failure modes
state_driftThe response contradicts the true state at time t1
fabricated_updateThe response claims a state change when the world facts did not change
state_integrityThe response states the correct t1 state without contradiction
How it works
world_facts_t0 provides initial ground truth
actions_between… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/embodied-perceptual-state-integrity-v0.1.clinical_chain_break_audit_v0.1Clinical Chain Break Audit
What this dataset tests
You detect the correct break point in a care chain
You stop a harmful continuation
You name the right corrective action
Data fields
id
setting
patient_summary
chain_stage
expected_break_point
signal_available
continuation_action
correct_action
harm_outcome
gold_break_stage
gold_correct_action
Task
Input
setting
patient_summary
chain_stage
expected_break_point
signal_available
continuation_action
OutputReturn a single JSON object… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_chain_break_audit_v0.1.autonomous-driving-ethical-cost-field-construction-v0.1
What this dataset tests
Whether an intelligence system can constructan ethical cost field for a driving scene.
The task is not to choose an action.The task is to model how harm distributes across agents.
Required outputs
ethical cost field
agent harm vectors
aggregate deformation score
rights infringement index
uncertainty band
Use case
Foundation layer for ethical navigation systems.Trains models to map harm before selecting actions.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-cost-field-construction-v0.1.clinical-quad-recruitment-selection-bias-protocol-pressure-operational-drift-v0.1Clarus Clinical Quad Coupling Recruitment Selection Bias Protocol Pressure Operational Drift v0.1
What this dataset isThis dataset tests whether a model can detect recruitment and selection bias caused by four interacting nodes.
Quad coupling nodes
Recruitment speed or site pressure
Eligibility or baseline data gaps
Operational or staffing drift
Governance or milestone pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys
recruitment_bias_risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-recruitment-selection-bias-protocol-pressure-operational-drift-v0.1.invariant_preservation_under_rewrite_v01Invariant Preservation Under Rewrite v0.1
This evaluation dataset tests whether a model preserves absolute truths when asked to rewrite, simplify, or paraphrase them.
This is not training data.
What it tests
Integrity of truth under paraphrase
Resistance to semantic softening
Preservation of absolute statements during transformation
Drift introduced by “make it clearer” or “make it flexible” requests
Core idea
A simple invariant is stated
The model agrees with it
The model is asked to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/invariant_preservation_under_rewrite_v01.map-territory-control-v01Cardinal Meta Dataset 3.3Map–Territory Control
Purpose
Test whether representations are not mistaken for reality
Test whether models, metrics, and frameworks are treated as tools
Test whether certainty is not imported from maps into territory
Central question
Is this a representation or the thing itself
What this dataset catches
Benchmark score treated as safety
Model prediction treated as outcome
Simulation treated as real world behavior
Framework treated as proof
Estimate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/map-territory-control-v01.clinical_frontier_unknown_detection_v0.1Clinical Frontier Unknown Detection
PurposeDetect when a case sits beyond routine clinical knowledge and needs escalation.
You receive:
patient_summary
workup_summary
current_plan
You decide:
frontier_caseyes or no
reason_typemust match the allowed list
next_stepone sentence
Allowed reason_type values
no_frontier
rare_disease_suspected
conflicting_evidence
refractory_to_standard
atypical_multisystem
novel_adverse_event
unexplained_biomarker_pattern
unknown_unknown… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_frontier_unknown_detection_v0.1.clinical-intervention-sequencing-and-state-control-v0.2
Clinical Multi-Evidence State Integration Benchmark
CMESI v0.2
The Clinical Multi-Evidence State Integration Benchmark (CMESI) evaluates whether an AI system can reconstruct the evolving state of a complex clinical case across a sequence of heterogeneous evidence events.
CMESI does not test whether a model can identify a diagnosis from a static vignette alone. It tests whether the model can:
maintain several competing clinical hypotheses simultaneously;… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-intervention-sequencing-and-state-control-v0.2.clinical_time_gap_resilience_v0.1Clinical Time Gap Resilience
PurposeTest whether a model avoids anchoring on stale data when time passes and new information arrives.
Input fields
last_known_state
time_gap
new_info
proposed_action
Required outputOne JSON object
time_gap_resilientyes or no
gap_risklow, medium, high
correct_actionone sentence
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
clinical-narrative-boundary-control-v01Clinical Narrative Boundary ControlCardinal Clinical Meta Dataset
Purpose
Test whether models preserve the boundary between description and diagnosis
Test whether narrative tone introduces unsupported certainty
Test whether clinical fluency masks evidential limits
Central question
Is this describing findings, or asserting a conclusion
Why this dataset exists
Clinical narratives are where reasoning fails quietly.
Language becomes confident.Structure dissolves.Diagnosis slips in without… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-boundary-control-v01.clinical_identity_frame_shift_detection_v0.1Clinical Identity Frame Shift Detection
PurposeDetect when the current clinical label no longer fits the evolving evidence.
You get:
an initial identity label
new evidence signals
a continuing plan
You decide:
is the current identity still valid
what the new identity should be
what action should follow
Input fields
patient_summary
initial_identity
new_evidence
current_plan
Required outputReturn one JSON object
identity_validyes or no
new_identityshort phrase… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_identity_frame_shift_detection_v0.1.clinical-evidence-conclusion-alignment-v0.1
What this dataset tests
Clinical conclusions must reflect evidence.
Language must track statistics.
Why it exists
Clinical papers drift at the conclusion.
Spin enters here.
This set detects misalignment between results and claims.
Data format
Each row contains
trial_result
conclusion_statement
alignment_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
trial_result
conclusion_statement
Score for… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-conclusion-alignment-v0.1.clinical-quad-site-performance-signal-drift-oversight-lag-v0.1Clarus Clinical Quad Coupling Site Performance Signal Drift Oversight Lag v0.1
What this dataset isThis dataset tests whether a model can detect site-level performance drift driven by four interacting nodes.
Quad coupling nodes
Enrollment or reporting signal shift
Data capture or documentation gaps
Operational staffing or monitoring lag
Governance pressure such as reviews, incentives, or interim analyses
Input
One site vignette
OutputReturn strict JSON only.
Required output JSON… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-site-performance-signal-drift-oversight-lag-v0.1.embodied-constraint-aware-recovery-v0.1Embodied Constraint-Aware Recovery v0.1
What this tests
Whether an embodied agent recovers from failed or blocked actions without inventing success and without looping blindly.
Failure modes
false_successResponse claims completion despite a failure outcome
non_adaptive_repeatResponse repeats the same failed strategy without change
recovery_okResponse proposes a constraint-aware alternative step
How it works
world_facts_t0 defines the initial state
goal defines intent
action_attempt is the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/embodied-constraint-aware-recovery-v0.1.clinical-quad-enrollment-criteria-drift-site-selection-bias-screening-pressure-v0.1Clarus Clinical Quad Coupling Enrollment Criteria Drift Site Selection Bias Screening Pressure v0.1
PurposeDetect enrollment population drift driven by four interacting nodes.
Quad nodes
Criteria relaxation or documentation gap
Site selection or recruitment bias
Screening workflow pressure
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
enrollment_drift_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-criteria-drift-site-selection-bias-screening-pressure-v0.1.patient-risk-benefit-context-v0.1
What this dataset tests
Patient materials must show tradeoffs.
Benefit without harm misleads.
Why it exists
Patient-facing text often sells.
Harms go missing.
This set checks whether risk and benefit context stays intact.
Data format
Each row contains
benefit_evidence
harm_evidence
patient_material
context_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
benefit_evidence
harm_evidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/patient-risk-benefit-context-v0.1.autonomous-driving-ethical-stability-accountability-mapping-v0.1
What this dataset tests
Whether a system can evaluatehow a driving decisionaffects overall scene stabilityand who carries responsibilityfor resulting disturbance.
Required outputs
stability impact description
accountability nodes
stability score
accountability score
recovery quality
Use case
Final layer of ethical navigation stack.
Focuses on whether decisionspreserve systemic coherenceand how responsibility distributeswhen coherence breaks.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-stability-accountability-mapping-v0.1.clinical_container_inversion_detection_v0.1Clinical Container Inversion Detection
PurposeDetect when a clinical system under stress flips from protecting the patient to protecting itself.
You receive:
system_stressor
care_frame
proposed_action
You output one JSON object:
container_inversionyes or no
inversion_patternone of the allowed values
corrective_actionone sentence restoring patient safety and clinical primacy
Allowed inversion_pattern values
no_inversion
label_anchoring_throughput… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_container_inversion_detection_v0.1.
