datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
system-stability-collapse-benchmark-casses-v0.1CASSES — Collapse Analysis in State-Space Evaluation Suite
Overview
CASSES is a diagnostic benchmark designed to test whether machine learning systems can detect instability and collapse in dynamic systems.
Most AI benchmarks evaluate models on tasks such as classification, language generation, or reasoning over static data.
CASSES evaluates a different capability:
state-space stability understanding.
The benchmark tests whether a model can identify when a system is approaching a collapse… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/system-stability-collapse-benchmark-casses-v0.1.epistemic_clarification_benchmark_v01
Purpose
Measure a model’s ability to detect when the question itself is flawed.
What this tests
contradiction detection
premise instability
ethical incoherence
context awareness
refusal clarity without moralizing
Format
Each row asks for:
the correct classification of the prompt
the expected response_target
a short reason_trace showing where the premise breaks
Why this matters
Modern LLMs fail not just by answering incorrectly but by… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/epistemic_clarification_benchmark_v01.latent-cross-coupling-instability-benchmark-v0.1Latent Cross Coupling Instability Benchmark v0.1
Overview
Some systems collapse not because visible signals indicate imminent failure, but because hidden interactions between subsystems amplify stress in ways that are not directly observable.
This benchmark evaluates whether machine learning systems can detect instability caused by latent cross-coupling interactions.
In these scenarios:
• subsystem A appears stable
• subsystem B appears stable
• observed coupling appears moderate
Yet hidden… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/latent-cross-coupling-instability-benchmark-v0.1.protein-folding-instability-trajectory-benchmark-v0.2Protein Folding Instability Trajectory Benchmark v0.2
Overview
This benchmark evaluates whether models can detect protein folding instability trajectories.
Unlike many protein AI tasks, the objective here is not to predict the final folded structure.
Instead the model must determine whether a folding trajectory is moving toward:
stable folding convergence
or
future misfold instability.
Protein folding occurs within an energy landscape containing multiple basins.
A folding process may converge… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/protein-folding-instability-trajectory-benchmark-v0.2.latent-coupling-instability-benchmark-v0.1Latent Coupling Instability Benchmark v0.1
Overview
This benchmark evaluates whether machine learning models can detect system collapse caused by interactions between variables rather than individual variables alone.
Many real-world systems fail not because a single factor becomes extreme, but because multiple factors interact in ways that amplify instability. These interaction effects are common in complex systems such as:
infrastructure networks with feedback loops
financial systems with… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/latent-coupling-instability-benchmark-v0.1.false-stability-collapse-benchmark-v0.1
False Stability Collapse Benchmark v0.1
Overview
This benchmark evaluates whether machine learning systems can detect future collapse when surface indicators still appear stable.
Many complex systems show calm behavior immediately before sudden failure.
These conditions are often referred to as false stability or metastable states.
Examples include:
financial markets before crashes
ecosystems approaching tipping points
infrastructure networks prior to cascading failures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/false-stability-collapse-benchmark-v0.1.trajectory-aliasing-instability-benchmark-v0.1Trajectory Aliasing Instability Benchmark v0.1
Overview
Some systems appear identical when observed at a single moment but evolve toward very different outcomes.
This phenomenon is called trajectory aliasing.
Two states may appear similar in surface variables yet belong to different underlying trajectories.
One trajectory leads toward stability.
The other leads toward collapse.
This benchmark evaluates whether machine learning systems can detect future collapse when trajectory signals are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/trajectory-aliasing-instability-benchmark-v0.1.recovery-window-instability-benchmark-v0.1Recovery Window Instability Benchmark v0.1
Overview
Some systems remain recoverable only for a limited period.
A system may not yet be collapsed, yet the window for successful intervention may already be narrowing or nearly closed.
This benchmark evaluates whether machine learning systems can detect recovery-window failure geometry.
In these scenarios, two systems may look similar in surface state and even similar in directional movement, but differ in one critical respect:
one still has a… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/recovery-window-instability-benchmark-v0.1.trajectory-divergence-instability-benchmark-v0.1
Trajectory Divergence Instability Benchmark v0.1
Overview
This benchmark evaluates whether machine learning systems can detect future collapse by interpreting directional system dynamics rather than static system state.
Many real-world systems appear stable when viewed at a single point in time.
However, subtle directional signals may indicate that the system is already moving toward instability.
This phenomenon occurs in many domains:
financial markets approaching… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/trajectory-divergence-instability-benchmark-v0.1.boundary-masking-instability-benchmark-v0.1Boundary Masking Instability Benchmark v0.1
Overview
Many systems fail not because instability is visible but because instability is masked by apparently stable surface indicators.
This benchmark evaluates whether machine learning systems can detect collapse risk when instability boundaries are hidden behind misleading surface signals.
Surface indicators may appear stable while the system is already close to a failure boundary.
This occurs in many domains
financial liquidity collapse… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/boundary-masking-instability-benchmark-v0.1.coupling-cascade-instability-benchmark-v0.1oupling Cascade Instability Benchmark v0.1
Overview
Some systems do not fail because one part becomes unstable in isolation.
They fail because subsystems that appear manageable on their own become dangerous when coupling between them amplifies local stress into global cascade failure.
This benchmark evaluates whether machine learning systems can detect coupling-driven cascade instability.
In these scenarios:
subsystem A may appear tolerable
subsystem B may appear tolerable
local stress may… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/coupling-cascade-instability-benchmark-v0.1.
