jub-aer/ConsistencyBench-interpretability
ConsistencyBench-Interpretability Extension White-box mechanistic analysis of logical inconsistency using Qwen/Qwen2.5-1.5B-Instruct (local, full activation access) as a dedicated interpretability testbed, distinct from the 17-model black-box leaderboard. Contents layer_probe_results.csv - per-layer logistic-regression probe accuracy for decoding "will this response be inconsistent?" directly from residual-stream activations activation_patching.csv - literal… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/ConsistencyBench-interpretability.
011
uploaded the files
updated readme
initial commit
