aaroncool9/kev-0.5b
Model card: kev-0.5b
kev-0.5b is a decision model. It takes one document (the state) and a set of typed questions, and returns a probability distribution for each question in one forward pass. It does not generate text.
It is a LoRA adapter plus a small pointer head on top of Qwen/Qwen2.5-0.5B. It reproduces the architecture that Archer Hume inferred for TypeSafe's Jev in *Jev's Architecture Unmasked*, and it serves TypeSafe's public /v1/systemone API contract.
This checkpoint is a research prototype trained on a laptop. It shows that the mechanism works. It is not a production model and it is not Jev.
- Hub: jaredpalmer/kev-0.5b (this repo, run
kev) - Code, training recipe, evaluation and demo: github.com/jaredpalmer/kev
- Weights: GitHub release `v0.1.0`,
kev-0.5b.tar.gz(38 MB; LoRA adapteradapter_model.safetensors, headhead.pt, tokenizer files,eval.json, training log). SHA-25615639f79…6e12f8, full digest in the sidecar.sha256. Extract toruns/kev/. Weights are not committed to git.
Model details
Intended use
Intended. Research on decision models: calibration of direct probability readouts, shared-state / isolated-question attention, option-order sensitivity, and API-level compatibility with TypeSafe's System One contract. Local demos and teaching.
Not intended. Any production decision that affects people: moderation, fraud, credit, hiring, medical or legal routing. The model's knowledge is limited to a 0.5B backbone, its calibration is only verified on the training distributions, and its outputs on unfamiliar tasks have not been measured.
How the model is used
Input is one packed token sequence:
<state> …state… <q> instr <opt> o1 </opt> <opt> o2 </opt> … <decide> <q> … <decide> …- The attention mask lets a question token see the state and its own branch only. Questions cannot see each other.
- Each branch restarts position ids after the state.
- For each question, the head scores every
</opt>hidden state against the<decide>hidden state and applies softmax. - Application code turns the distributions into the API answer:
choice/confidencefor Choice,p(yes)for Noul, expected level for Score.
Reserved tokens are existing Qwen special tokens (<|fim_prefix|>, <|fim_middle|>, <|box_start|>, <|box_end|>, <|fim_suffix|>). User text is sanitized so it cannot produce them.
Serve with python -m kev.serve --run runs/kev and call POST /v1/systemone, or use typesafe-sdk with base_url="http://127.0.0.1:8009".
Training data
Six public datasets, converted to TypeSafe-shaped requests and rendered with the same code path used at serving time (api.to_record()). 1,500 records were sampled per source from the standard train splits, giving 9,000 records and 13,500 questions (4,500 Choice, 6,000 Noul, 3,000 Score).
Rendering variation applied at conversion time: ~30% null option descriptions, ~10% structured {"what": …} descriptions, ~15% structured {"question", "focus"} instructions, ~32% states wrapped as objects or arrays ({"document"}, {"ticket": {"channel","body"}}, [{"role","content"}]).
Augmentation applied once per record before encoding: option order shuffled; with probability 0.10 the true option replaced by other: None of the above; with probability 0.15 an irrelevant distractor option added.
No LLM-generated data. No human annotation beyond the original datasets.
Training procedure
This checkpoint predates two loss terms that are now defaults in kev/train.py: the ordinal term for Score (--ord_w) and the permutation-consistency KL for Choice (--perm_kl). To reproduce this checkpoint exactly:
uv run python -m kev.train --n_per_source 1500 --epochs 2 --accum 8 --perm_kl 0 --ord_w 0 --out runs/kevNote that augmentation is now re-applied every epoch rather than fixed at encode time, so a re-run will not be bit-identical.
Evaluation
Held-out test / validation splits of the same six sources, 150 records per source, 1,350 questions, seed 1. Full results in runs/kev/eval.json.
Accuracy and calibration
Baselines: Qwen/Qwen2.5-0.5B (raw) and Qwen/Qwen2.5-0.5B-Instruct (chat template), same rendered text, next-token logits over option letters A–H; not run for K = 77. ECE uses 10 equal-width bins on the top probability.
Temperature scaling
Fit on even-indexed records, tested on odd-indexed: T = 1.47. Held-out NLL 0.505 → 0.481, ECE 0.057 → 0.031. The model is mildly over-confident before scaling.
Mechanism tests
Limitations
- In-distribution only. All numbers above are on held-out splits of the training datasets. Out-of-source generalization has not been measured for this checkpoint.
- Small backbone. 0.5B parameters. On the TypeSafe docs' structured-criteria example the model picks
return_policywhere Jev picksreturn_status. Reading comprehension (BoolQ 0.75, MNLI 0.75) is far below state of the art. - Narrow task coverage. Six datasets and about ten instruction templates. Code, tables, multi-turn chat, arithmetic, and multi-step conditions are untrained.
- Order sensitivity remains. 7% argmax flips and a p90 probability spread of 0.25 under option reordering. A threshold near a decision boundary can change the action.
- Score confidence is a stand-in.
1 − E|level − mode| / (L − 1); TypeSafe's formula is unpublished. - Calibration is not a guarantee. ECE 0.03 after temperature scaling on these sources says nothing about calibration on a new workflow. Proper scoring rules give the right incentive; they do not remove the need for outcome data.
- Inherited limitations from Qwen2.5-0.5B and from the datasets, including their label noise, demographic skews (e.g. Yelp, banking intents), and English-only coverage.
Bias, risks and recommendations
The training sets carry the biases of their sources: US-centric news categories, English banking terminology, restaurant reviews, and crowd-sourced NLI labels. The model will mirror them.
Direct probability outputs look authoritative. A confidence: 0.92 from this model is a statistic about its own distribution over three options, not a verified probability of being right. Do not threshold on it for consequential decisions without measuring calibration on your own labelled outcomes first.
The question-isolation property is a real safety feature (one question's text cannot manipulate another's answer) and was verified. The delimiter-forgery protection was verified for the five reserved tokens. Other prompt-injection routes through the state text have not been studied.
Environmental impact
One training run: ~1.75 h on a single Apple M5 laptop SoC at roughly 30–40 W, i.e. about 0.06 kWh. Evaluation and smoke runs add a similar amount. This is small.
Citation
@software{kev2026,
title = {kev: a laptop-scale reconstruction of a Jev-style decision model},
author = {Palmer, Jared},
year = {2026},
url = {https://github.com/jaredpalmer/kev}
}
@misc{hume2026jev,
title = {Jev's Architecture Unmasked},
author = {Hume, Archer},
year = {2026},
url = {https://archerhume.com/posts/jevs-architecture-unmasked}
}Contact
Open an issue at github.com/jaredpalmer/kev.
