clarusc64-benchmark
clinical-stability-benchmark
Benchmark Documentation
Core
benchmark_structure.md
benchmark_matrix.md
datasets.md
Evaluation
evaluation_framework.md
transfer_matrix.md
clarus_score.md
Robustness
missing_data_protocol.md
imbalance_protocol.md
robustness_suite.md
Theory
stability_manifold.md
stability_topology.md
stability_mechanisms.md
Results
baseline_results.md
leaderboard.md
Clarus Clinical Stability Benchmark
The Clarus Clinical Stability Benchmark evaluates whether machine learning… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-stability-benchmark.eval-trap-stability-manifold-benchmark-v0.2
Eval Trap Stability Manifold Benchmark v0.2
This repository provides a synthetic benchmark for testing whether models can distinguish between content confidence and system viability.
The benchmark is built to expose the evaluation trap:
A model assigns high confidence to a proposed configuration even though the system executing that configuration is mathematically unstable.
Core idea
Most predictive systems optimize for content accuracy.
This benchmark tests something… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/eval-trap-stability-manifold-benchmark-v0.2.system-stability-collapse-benchmark-casses-v0.1CASSES — Collapse Analysis in State-Space Evaluation Suite
Overview
CASSES is a diagnostic benchmark designed to test whether machine learning systems can detect instability and collapse in dynamic systems.
Most AI benchmarks evaluate models on tasks such as classification, language generation, or reasoning over static data.
CASSES evaluates a different capability:
state-space stability understanding.
The benchmark tests whether a model can identify when a system is approaching a collapse… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/system-stability-collapse-benchmark-casses-v0.1.epistemic_clarification_benchmark_v01
Purpose
Measure a model’s ability to detect when the question itself is flawed.
What this tests
contradiction detection
premise instability
ethical incoherence
context awareness
refusal clarity without moralizing
Format
Each row asks for:
the correct classification of the prompt
the expected response_target
a short reason_trace showing where the premise breaks
Why this matters
Modern LLMs fail not just by answering incorrectly but by… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/epistemic_clarification_benchmark_v01.latent-cross-coupling-instability-benchmark-v0.1Latent Cross Coupling Instability Benchmark v0.1
Overview
Some systems collapse not because visible signals indicate imminent failure, but because hidden interactions between subsystems amplify stress in ways that are not directly observable.
This benchmark evaluates whether machine learning systems can detect instability caused by latent cross-coupling interactions.
In these scenarios:
• subsystem A appears stable
• subsystem B appears stable
• observed coupling appears moderate
Yet hidden… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/latent-cross-coupling-instability-benchmark-v0.1.protein-folding-instability-trajectory-benchmark-v0.2Protein Folding Instability Trajectory Benchmark v0.2
Overview
This benchmark evaluates whether models can detect protein folding instability trajectories.
Unlike many protein AI tasks, the objective here is not to predict the final folded structure.
Instead the model must determine whether a folding trajectory is moving toward:
stable folding convergence
or
future misfold instability.
Protein folding occurs within an energy landscape containing multiple basins.
A folding process may converge… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/protein-folding-instability-trajectory-benchmark-v0.2.
