datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maritime-bunker-consumption-voyage-plan-coherence-risk-v0.1What this repo is for
Detect when fuel burn stops matching voyage plan.
You use it to flag:
unexpected efficiency loss
reserve margin collapse
speed pushing fuel beyond plan
weather masking burn drift
Why it matters
Fuel is the largest variable cost in shipping
ai-5node-key-buf-lag-cpl-secret-leak-v0.1
What this repo does
This dataset models secret leakage cascades in AI agent operations. It detects when secret exposure risk rises, protective buffers weaken, governance lag delays revoke and purge actions, and tight coupling through shared logs, tickets, and tool chains crosses the five-node cascade threshold into an unrecoverable secret leakage cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The fifth node… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-key-buf-lag-cpl-secret-leak-v0.1.maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
reasoning-trajectory-stability-controls-v0.1
Reasoning Trajectory Stability Controls v0.1
A SIOS research dataset for detecting whether a reasoning trajectory remains structurally stable, identifying the control introduced into the trajectory, locating where that control first becomes operationally visible, and determining whether the control succeeds or fails.
Repository:
ClarusC64/reasoning-trajectory-stability-controls-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-trajectory-stability-controls-v0.1.legal-disclosure-coherence-breach-detection-v0.1What this dataset is
You receive
disclosure duty
material
timing
defence access
prejudice signals
You decide
Does disclosure behaviour match the legal duty
Answer
coherent
or
incoherent
Why this matters
Many unsafe convictions arise from disclosure failure.
This dataset measures the structural gap between duty and behaviour.
clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Integrity v0.1
PurposeDetect adjudication drift driven by four interacting nodes.
Quad nodes
Endpoint cluster shift
Blinding gap or reviewer dominance
Operational or vendor process change
Governance submission or review pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
adjudication_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1.cascade-multi-ai-redteam-miss-exploit-patch-trust-collapse-v0.1
What this repo does
This dataset tests whether a model can detect a security cascade in AI deployment.
You provide structured signals about:
red-team coverage and disclosure
exploitability and incident rate
patch latency and rollout friction
downstream dependency depth
trust decay and regulatory attention
The model predicts whether the interaction crosses into a cascade event.
Core quad
The structural quad inside this cascade:
red_team_coverage
exploitability_index… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-multi-ai-redteam-miss-exploit-patch-trust-collapse-v0.1.clinical_narrative_negative_evidence_handling_v0.4Clinical Narrative Negative Evidence Handling v0.4
Purpose
Test whether a model handles negative evidence without narrative spin.
This version adds
timeline steps
cross trial negative carryover
suppression pressure prompts
explicit evidence status and submission positioning
Input columns
data_anchor
negative_pressures
draft_narrative
audience
timeline_step
Model task
Return one JSON object
negative_flagslist of short labels
evidence_statusexploratory, mixed, negative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_narrative_negative_evidence_handling_v0.4.maritime-bunker-consumption-route-coherence-risk-v0.1What this repo is for
Detect when voyage routing decisions push fuel burn out of plan.
You use it to flag
schedule recovery that triggers fuel spikes
efficiency loss hidden behind stable routing
low bunker margin with weak next-port options
compound risk from route drift plus speed-up
Why it matters
Fuel drift turns into cost spikes fast.
It also drives unplanned bunkering and schedule instability.
clinical-drv-atlas-perturbation-response-stability-mapping-v0.1What this dataset tests
Whether a model can classify response topologyafter a controlled perturbation.
It rewards
correct topology
recognition of cross-system coupling
recovery timing
Response topologies
rapid_return
delayed_recovery
overshoot_instability
oscillatory_instability
collapse
Typical failures
confusing overshoot with oscillation
ignoring coupling direction
calling delayed recovery stable
Suggested prompt wrapper
System
You map perturbation response… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-perturbation-response-stability-mapping-v0.1.clinical-quad-unblinding-sae-cluster-media-leak-trial-halt-decision-v0.1Clinical Quad Unblinding SAE Cluster Media Leak Trial Halt Decision v0.1
Each row is a site weekly snapshot.
Core quad
Emergency unblindingSAE clusterMedia leak riskTrial halt decision risk
Target
label_trial_halt_risk_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6
What this repo does
This repository contains a Clarus v0.6 intervention pathway dataset focused on respiratory collapse dynamics.
The dataset evaluates whether a model can determine if a proposed intervention meaningfully stabilizes a deteriorating respiratory system.
The task requires reasoning from:
system state
trajectory toward instability
boundary geometry
recovery geometry
intervention vector
projected trajectory consequence
The model cannot read the answer directly.
It must… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6.clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1What this dataset tests
Whether a model can extract the minimal cross-system coherence factor setthat explains resilience or vulnerability.
It rewards
minimal factor selection
correct coupling recognition
ranking by dominance
Coherence factor labels
buffering_capacity_high
buffering_capacity_low
variance_damping_high
variance_damping_low
autonomic_inflammatory_coupling
sleep_metabolic_coupling
stress_inflammation_coupling
immune_metabolic_instability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1.clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0
ClarusC64/clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0
What this repo does
This repository provides a Clarus v1.0 benchmark for postoperative collapse under a four-variable clinical quad:
surgical_stress
buffer_capacity
lag_burden
coupling_stress
The v1.0 upgrade is Closed-Loop Control Geometry.
The task is no longer limited to detecting deterioration or ranking one intervention against another.
It tests whether a controller can:
choose the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0.acquisition-plausibility-integrity-medimg-v01Acquisition Plausibility Integrity v01
What this dataset is
This dataset evaluates whether a system can judge if a claimed imaging outcome is physically or technically possible given the modality and acquisition parameters.
You give the model:
An imaging modality and protocol
Acquisition parameters
A claimed diagnostic capability
You ask one question.
Can this scan
contain this information
at all
Why this matters
Medical imaging errors often begin before interpretation.
Common failure… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/acquisition-plausibility-integrity-medimg-v01.aerogel-structural-manifold-integrity-v0.1Goal
Detect when an aerogel loses structural integrity before visible collapse.
Core idea
Aerogel failure is not a single crack.It is a distortion of the vibration–density–pore manifold.
Three signals must stay coherent:
densityelastic moduluspore network structure
When they decouple, collapse follows.
Inputs
bulk density
nanoindentation modulus
pore size distribution
load cycling
acoustic or strain indicators
Required outputs
structural_coherence_score
manifold_distortion_rate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aerogel-structural-manifold-integrity-v0.1.legal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does
You receive
doc description
date
author
recipients
privilege basis
redaction choice
context
waiver flags
You decide
coherent
or
incoherent
Daily use
privilege log QC
waiver risk detection
disclosure challenge prep
ai-carbon-claim-metering-coherence-breach-v0.1What this repo is for
Detect when carbon claims
diverge from metered reality.
Flags
net-zero claims with high load on high-intensity grids
weak market instruments vs actual consumption
annual matching used to hide high-carbon hours
strong claims without strong time matching
clinical-prescription-pharmacy-dispense-coherence-risk-v0.1What this repo is for
Detect when
a prescription exists
but pharmacy dispense
does not happen in time
Common breaks
stockout
verification delay
clarification needed
queue delay for discharge meds
Examples you can use
urgent anticoagulant delayed
antibiotic not dispensed due to stockout
TTO delayed so discharge stalls
You use it to flag
missed dose risk
discharge delay risk
reasoning-drift-onset-detection-v0.2A SIOS structured reasoning-state benchmark for detecting when a reasoning trajectory loses a governing constraint, identifying the structural form of that drift, and assessing whether the failure is repaired.
Repository:
ClarusC64/reasoning-drift-onset-detection-v0.2
Version:
0.2.0
Publisher:
Clarus Invariant
Framework:
SIOS
Benchmark identity
Reasoning Drift Onset Detection v0.2 is not a single-label classification benchmark.
It is a structured reasoning-state benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-drift-onset-detection-v0.2.legal-time-entry-billing-narrative-scope-coherence-risk-v0.1What this dataset does
You receive
scope
billing guidelines
time entries
fee earner level
billing narrative
duration and rates
flags
You decide
coherent
or
incoherent
Daily use
fee dispute risk scan
scope drift scan
block billing detection
seniority mismatch detection
clinical-meaning-integration-fragmentation-analysis-v0.1What this dataset tests
Whether an intelligence system can evaluatethe integrity of a patient’s meaning-making processduring illness and recovery.
Required outputs
coherence integrity score
fragmentation markers
denial indicators
adaptive reframing presence
narrative stability index
meaning failure mode
Use case
Second layer of the Healing Narrative Coherence Corpus.
autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1What this dataset tests
Whether a system can score traffic-flow coherence
before and after an ego action.
This is not collision detection.
It measures systemic stability.
Required outputs
pre_action_coherence_score
post_action_coherence_score
coherence_delta
shockwave_generation_flag
braking_propagation_depth
systemic_risk_score
Scoring conventions
coherence scores range 0 to 1
coherence_delta may be negative or positive
shockwave flag is 0 or 1
braking propagation depth… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1.reasoning-constraint-loss-attribution-v0.1
Reasoning Constraint Loss Attribution v0.1
A SIOS research dataset for identifying when a governing constraint ceases to regulate a reasoning trajectory, locating the first point of loss, attributing the lost constraint, and identifying the structural mechanism that produced the loss.
Repository:
ClarusC64/reasoning-constraint-loss-attribution-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity
Reasoning Constraint Loss Attribution… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-constraint-loss-attribution-v0.1.clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1
Goal
Test if a model can hold separate reasoning streams at once
Detect constraint dismissal
Detect bleed-over where one stream turns into claims in the other
What it measures
streams_heldResponse acknowledges and maintains both streams
bleed_overConstraint stream improperly becomes a medical claim, or vice versa
premature_synthesisResponse forces a single solution that silences one stream
assumption_collapseResponse drops a premise entirely
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.counterfactual-action-outcome-fidelity-v0.1
What this dataset tests
Whether claimed outcomes in counterfactual scenariosfollow from the proposed actionunder the stated condition.
It checks causality and constraints.
Why this exists
Models often give outcomes that sound plausible.
They ignore:
physics
logical dependency
hard limits
safety constraints
This dataset flags that.
Data format
Each example includes:
counterfactual condition
proposed action
claimed outcome
causal check
constraint… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/counterfactual-action-outcome-fidelity-v0.1.clinical-counterfactual-branch-construction-v0.1What this dataset tests
Whether a model can generate clinically plausible counterfactual decision pointsanchored to a real patient timeline.
Required outputs
decision_point_id
counterfactual_action
plausibility_score_0_100
branching_constraints
Counterfactual types
timing_shift
choice_substitution
disposition_change
intensity_change
Typical failures
proposing actions not linked to the real timeline
omitting safety constraints
producing fantasy improvements without… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-counterfactual-branch-construction-v0.1.clinical-perception-intervention-justification-v0.1Clinical Perception–Intervention Justification v0.1
Goal
Test whether actions follow directly from perceptual evidence
Detect interventions that appear without a visual cause
Detect escalation that exceeds image-supported severity
What it measures
action_without_causeAn intervention is proposed with no supporting image evidence
over_escalationThe action exceeds what the visual severity supports
justification_okThe response links perception to action explicitly or proportionally
How it… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-perception-intervention-justification-v0.1.legal-billing-narrative-time-entry-coherence-risk-v0.1What this dataset does
You receive
time entries summary
phase and codes
invoice narrative
totals
dup flags
client updates
You decide
coherent
or
incoherent
Daily use
bill narrative QC
time entry duplication detection
dispute risk flagging
clinical-imaging-report-action-coherence-risk-v0.1What this repo is for
Detect when
imaging answers a question
but the system fails to
acknowledge and act
Common breaks
critical report not acknowledged
report late for urgent indication
action delayed after critical finding
negative result not integrated
Examples you can use
CTPA positive but anticoag not started
CT head bleed not escalated
CT perforation with delayed surgery
You use it to flag
missed diagnosis risk
treatment delay risk
