bdougie/smollm3-pokemon-red-lora
017
smollm3-pokemon-red-lora
One adapter for every seat of the pokemon-kafka expedition crew (the seat is the system prompt): the Wheelman's battle calls, the Extractor's puzzle menu choices, the Forger's reading of bodies and gate sentences, the Narrator's play-by-play, the operator handoffs.
- Base
HuggingFaceTB/SmolLM3-3B, bf16 LoRA r32 on q,k,v,o,gate,up,down; trained 2026-09-07 with empirical-evidence'sautotune.train_sft, 900 steps over 25,053 rows (about 0.6 epoch, 32 min on an RTX 5090). - Data: bdougie/pokemon-red-sft. Every row is measured from a run of Pokémon Red on the cartridge: the sprite table, the screen, the bag, the battle struct. No recalled game facts.
Held-out gate (2,783 rows; scored domains below; tuned vs base SmolLM3-3B)
For the Forger's own domains the dedicated adapter bdougie/smollm3-pokemon-forger-lora scores higher (body 0.82, outcome 0.66). Full numbers in eval.json; the training log in train.log.
Source code
- Trained and gated by pcc-labs/empirical-evidence (
autotune.train_sft,autotune.eval_heldout; packaging inscripts/package_lora.sh) on the dataset bdougie/pokemon-red-sft. - The crew it is a seat in: pcc-labs/pokemon-kafka. The full record is the white paper Training an Archetype and docs/forger-adapter.md.
