datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maritime-bunker-consumption-voyage-plan-coherence-risk-v0.1What this repo is for
Detect when fuel burn stops matching voyage plan.
You use it to flag:
unexpected efficiency loss
reserve margin collapse
speed pushing fuel beyond plan
weather masking burn drift
Why it matters
Fuel is the largest variable cost in shipping
ai-5node-key-buf-lag-cpl-secret-leak-v0.1
What this repo does
This dataset models secret leakage cascades in AI agent operations. It detects when secret exposure risk rises, protective buffers weaken, governance lag delays revoke and purge actions, and tight coupling through shared logs, tickets, and tool chains crosses the five-node cascade threshold into an unrecoverable secret leakage cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The fifth node… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-key-buf-lag-cpl-secret-leak-v0.1.clinical-stability-benchmark
Benchmark Documentation
Core
benchmark_structure.md
benchmark_matrix.md
datasets.md
Evaluation
evaluation_framework.md
transfer_matrix.md
clarus_score.md
Robustness
missing_data_protocol.md
imbalance_protocol.md
robustness_suite.md
Theory
stability_manifold.md
stability_topology.md
stability_mechanisms.md
Results
baseline_results.md
leaderboard.md
Clarus Clinical Stability Benchmark
The Clarus Clinical Stability Benchmark evaluates whether machine learning… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-stability-benchmark.reasoning-trajectory-stability-controls-v0.1
Reasoning Trajectory Stability Controls v0.1
A SIOS research dataset for detecting whether a reasoning trajectory remains structurally stable, identifying the control introduced into the trajectory, locating where that control first becomes operationally visible, and determining whether the control succeeds or fails.
Repository:
ClarusC64/reasoning-trajectory-stability-controls-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-trajectory-stability-controls-v0.1.airframe-sios-hidden-geometry
Airframe SIOS Hidden Geometry Benchmark
Overview
Airframe SIOS Hidden Geometry is a synthetic relational-reasoning benchmark designed to test whether a system can recover a globally coherent labelled graph when local observations are incomplete, overlapping, conflicting, or actively misleading.
The benchmark is built around two central distinctions:
Metric Evidence ≠ Relational Structure
Local Plausibility ≠ Global Coherence
Each example contains four labelled… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/airframe-sios-hidden-geometry.legal-disclosure-coherence-breach-detection-v0.1What this dataset is
You receive
disclosure duty
material
timing
defence access
prejudice signals
You decide
Does disclosure behaviour match the legal duty
Answer
coherent
or
incoherent
Why this matters
Many unsafe convictions arise from disclosure failure.
This dataset measures the structural gap between duty and behaviour.
reasoning-drift-onset-detection-v0.1
Important Evaluation Limitation
Version 0.1 uses a highly regular trajectory structure in which the first drift step is frequently located at Step 4 and visible failure commonly appears at Step 5.
This creates a positional shortcut: a model may achieve inflated onset-detection performance by learning the dataset construction pattern rather than analysing the reasoning trajectory.
Version 0.1 should therefore be treated as a task-definition and scorer-validation release, not as a… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-drift-onset-detection-v0.1.clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Integrity v0.1
PurposeDetect adjudication drift driven by four interacting nodes.
Quad nodes
Endpoint cluster shift
Blinding gap or reviewer dominance
Operational or vendor process change
Governance submission or review pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
adjudication_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1.reasoning-drift-onset-detection-v0.2A SIOS structured reasoning-state benchmark for detecting when a reasoning trajectory loses a governing constraint, identifying the structural form of that drift, and assessing whether the failure is repaired.
Repository:
ClarusC64/reasoning-drift-onset-detection-v0.2
Version:
0.2.0
Publisher:
Clarus Invariant
Framework:
SIOS
Benchmark identity
Reasoning Drift Onset Detection v0.2 is not a single-label classification benchmark.
It is a structured reasoning-state benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-drift-onset-detection-v0.2.maritime-bunker-consumption-route-coherence-risk-v0.1What this repo is for
Detect when voyage routing decisions push fuel burn out of plan.
You use it to flag
schedule recovery that triggers fuel spikes
efficiency loss hidden behind stable routing
low bunker margin with weak next-port options
compound risk from route drift plus speed-up
Why it matters
Fuel drift turns into cost spikes fast.
It also drives unplanned bunkering and schedule instability.
cascade-multi-ai-redteam-miss-exploit-patch-trust-collapse-v0.1
What this repo does
This dataset tests whether a model can detect a security cascade in AI deployment.
You provide structured signals about:
red-team coverage and disclosure
exploitability and incident rate
patch latency and rollout friction
downstream dependency depth
trust decay and regulatory attention
The model predicts whether the interaction crosses into a cascade event.
Core quad
The structural quad inside this cascade:
red_team_coverage
exploitability_index… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-multi-ai-redteam-miss-exploit-patch-trust-collapse-v0.1.clinical_narrative_negative_evidence_handling_v0.4Clinical Narrative Negative Evidence Handling v0.4
Purpose
Test whether a model handles negative evidence without narrative spin.
This version adds
timeline steps
cross trial negative carryover
suppression pressure prompts
explicit evidence status and submission positioning
Input columns
data_anchor
negative_pressures
draft_narrative
audience
timeline_step
Model task
Return one JSON object
negative_flagslist of short labels
evidence_statusexploratory, mixed, negative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_narrative_negative_evidence_handling_v0.4.euv-stochastic-yield-failure-horizon-routing-v0.1
Purpose
Yield does not collapse instantly.It drifts.
The early signal is coherence decay between:
source noiseresist responseLER spreaddefect clusteringyield loss
This dataset predicts:
how far the system driftedhow many lots remain before yield collapsewhat intervention stabilizes the line
Task
Return JSON:
drift_scorefailure_horizon_lotsintervention_route
Example
{"drift_score":0.58,"failure_horizon_lots":60,"intervention_route":"laser_tuning"}
Route… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/euv-stochastic-yield-failure-horizon-routing-v0.1.clinical-quad-unblinding-sae-cluster-media-leak-trial-halt-decision-v0.1Clinical Quad Unblinding SAE Cluster Media Leak Trial Halt Decision v0.1
Each row is a site weekly snapshot.
Core quad
Emergency unblindingSAE clusterMedia leak riskTrial halt decision risk
Target
label_trial_halt_risk_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6
What this repo does
This repository contains a Clarus v0.6 intervention pathway dataset focused on respiratory collapse dynamics.
The dataset evaluates whether a model can determine if a proposed intervention meaningfully stabilizes a deteriorating respiratory system.
The task requires reasoning from:
system state
trajectory toward instability
boundary geometry
recovery geometry
intervention vector
projected trajectory consequence
The model cannot read the answer directly.
It must… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6.clinical-drv-atlas-perturbation-response-stability-mapping-v0.1What this dataset tests
Whether a model can classify response topologyafter a controlled perturbation.
It rewards
correct topology
recognition of cross-system coupling
recovery timing
Response topologies
rapid_return
delayed_recovery
overshoot_instability
oscillatory_instability
collapse
Typical failures
confusing overshoot with oscillation
ignoring coupling direction
calling delayed recovery stable
Suggested prompt wrapper
System
You map perturbation response… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-perturbation-response-stability-mapping-v0.1.clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1What this dataset tests
Whether a model can extract the minimal cross-system coherence factor setthat explains resilience or vulnerability.
It rewards
minimal factor selection
correct coupling recognition
ranking by dominance
Coherence factor labels
buffering_capacity_high
buffering_capacity_low
variance_damping_high
variance_damping_low
autonomic_inflammatory_coupling
sleep_metabolic_coupling
stress_inflammation_coupling
immune_metabolic_instability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1.premier-league-game-state-coherence-mapping-v0.1What this dataset tests
Whether an intelligence system can scoreteam coherence in live match state.
Required outputs
team coherence score
lane control
compactness
pressure synchrony
progression coherence
coherence drop trigger
coherence zone map
Use case
First layer of Game State Coherence & Fragility Maps.
aerogel-structural-manifold-integrity-v0.1Goal
Detect when an aerogel loses structural integrity before visible collapse.
Core idea
Aerogel failure is not a single crack.It is a distortion of the vibration–density–pore manifold.
Three signals must stay coherent:
densityelastic moduluspore network structure
When they decouple, collapse follows.
Inputs
bulk density
nanoindentation modulus
pore size distribution
load cycling
acoustic or strain indicators
Required outputs
structural_coherence_score
manifold_distortion_rate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aerogel-structural-manifold-integrity-v0.1.ai-carbon-claim-metering-coherence-breach-v0.1What this repo is for
Detect when carbon claims
diverge from metered reality.
Flags
net-zero claims with high load on high-intensity grids
weak market instruments vs actual consumption
annual matching used to hide high-carbon hours
strong claims without strong time matching
clinical-prescription-pharmacy-dispense-coherence-risk-v0.1What this repo is for
Detect when
a prescription exists
but pharmacy dispense
does not happen in time
Common breaks
stockout
verification delay
clarification needed
queue delay for discharge meds
Examples you can use
urgent anticoagulant delayed
antibiotic not dispensed due to stockout
TTO delayed so discharge stalls
You use it to flag
missed dose risk
discharge delay risk
clinical-evidence-state-transition-fidelity-v0.1
Clinical Evidence State Transition Fidelity v0.1
A synthetic clinical reasoning dataset for evaluating whether an AI system can update a structured clinical state selectively, proportionately, and consistently when new evidence arrives.
Repository:
ClarusC64/clinical-evidence-state-transition-fidelity-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity
Clinical Evidence State Transition Fidelity v0.1 evaluates whether a model can… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-state-transition-fidelity-v0.1.acquisition-plausibility-integrity-medimg-v01Acquisition Plausibility Integrity v01
What this dataset is
This dataset evaluates whether a system can judge if a claimed imaging outcome is physically or technically possible given the modality and acquisition parameters.
You give the model:
An imaging modality and protocol
Acquisition parameters
A claimed diagnostic capability
You ask one question.
Can this scan
contain this information
at all
Why this matters
Medical imaging errors often begin before interpretation.
Common failure… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/acquisition-plausibility-integrity-medimg-v01.legal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does
You receive
doc description
date
author
recipients
privilege basis
redaction choice
context
waiver flags
You decide
coherent
or
incoherent
Daily use
privilege log QC
waiver risk detection
disclosure challenge prep
maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
legal-time-entry-billing-narrative-scope-coherence-risk-v0.1What this dataset does
You receive
scope
billing guidelines
time entries
fee earner level
billing narrative
duration and rates
flags
You decide
coherent
or
incoherent
Daily use
fee dispute risk scan
scope drift scan
block billing detection
seniority mismatch detection
clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0
ClarusC64/clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0
What this repo does
This repository provides a Clarus v1.0 benchmark for postoperative collapse under a four-variable clinical quad:
surgical_stress
buffer_capacity
lag_burden
coupling_stress
The v1.0 upgrade is Closed-Loop Control Geometry.
The task is no longer limited to detecting deterioration or ranking one intervention against another.
It tests whether a controller can:
choose the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0.eval-trap-stability-manifold-benchmark-v0.2
Eval Trap Stability Manifold Benchmark v0.2
This repository provides a synthetic benchmark for testing whether models can distinguish between content confidence and system viability.
The benchmark is built to expose the evaluation trap:
A model assigns high confidence to a proposed configuration even though the system executing that configuration is mathematically unstable.
Core idea
Most predictive systems optimize for content accuracy.
This benchmark tests something… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/eval-trap-stability-manifold-benchmark-v0.2.quantum-gate-sequence-instability-v0.1
quantum-gate-sequence-instability-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in quantum gate sequences.
Each row represents a simplified quantum circuit execution scenario described through observable device and circuit proxies.
The task is to determine whether the gate sequence remains executable inside a stable coherence window or becomes unstable.
Core stability idea
Quantum gate sequences become unstable when… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-gate-sequence-instability-v0.1.reasoning-constraint-loss-attribution-v0.1
Reasoning Constraint Loss Attribution v0.1
A SIOS research dataset for identifying when a governing constraint ceases to regulate a reasoning trajectory, locating the first point of loss, attributing the lost constraint, and identifying the structural mechanism that produced the loss.
Repository:
ClarusC64/reasoning-constraint-loss-attribution-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity
Reasoning Constraint Loss Attribution… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-constraint-loss-attribution-v0.1.
