ClarusC64/eval-trap-stability-manifold-benchmark-v0.2
Eval Trap Stability Manifold Benchmark v0.2 This repository provides a synthetic benchmark for testing whether models can distinguish between content confidence and system viability. The benchmark is built to expose the evaluation trap: A model assigns high confidence to a proposed configuration even though the system executing that configuration is mathematically unstable. Core idea Most predictive systems optimize for content accuracy. This benchmark tests… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/eval-trap-stability-manifold-benchmark-v0.2.
038
