LocalLLaMA/typed-decisions
Typed Decisions A benchmark for typed probabilistic decisions over shared state. You give a model one piece of unstructured state. It answers several typed questions about that state at once, and every answer is a probability distribution rather than a single label. The schema follows the System One primitives used by TypeSafe AI: noul, choice and score. A row replays against any API that implements that shape. This benchmark is independent. It is not affiliated with TypeSafe… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/typed-decisions.
Replace the estimated System One row with a measured Jev 1.13.0 result
card: point usage examples at the canonical org path
card: explain each reference row, add ceilings and a saturation estimate
card: correct the prior baseline, which had been fitted on the benchmark
card: add Adaptive Classifier specialist baselines (MiniLM and ModernBERT)
rewrite dataset card
remove build-time audit from the published dataset
full dataset: 400-case benchmark (test) + 1200-case training split, verified disjoint
card: specialist vs generalist evaluation modes
benchmark-only: single test split of 400; add measured score reference points
card: remove personal attribution from citation
typed-decisions: 400 cases, 2000 typed decisions across 4 workflows
initial commit
