CoolFace
Modelpublic

bdougie/smollm3-pokemon-red-lora

sourceHugging Faceupdated 19d agoView on Hugging Face
0likes17downloads
Model Card

smollm3-pokemon-red-lora

One adapter for every seat of the pokemon-kafka expedition crew (the seat is the system prompt): the Wheelman's battle calls, the Extractor's puzzle menu choices, the Forger's reading of bodies and gate sentences, the Narrator's play-by-play, the operator handoffs.

  • —Base HuggingFaceTB/SmolLM3-3B, bf16 LoRA r32 on q,k,v,o,gate,up,down; trained 2026-09-07 with empirical-evidence's autotune.train_sft, 900 steps over 25,053 rows (about 0.6 epoch, 32 min on an RTX 5090).
  • —Data: bdougie/pokemon-red-sft. Every row is measured from a run of Pokémon Red on the cartridge: the sprite table, the screen, the bag, the battle struct. No recalled game facts.

Held-out gate (2,783 rows; scored domains below; tuned vs base SmolLM3-3B)

metricrowsbasetunedmajority
battle-outcome3490.520.99
move-choice10610.000.95
npc-dialogue/body1320.210.77
npc-dialogue/outcome1320.420.630.61 (always "talk")
gate-text/gate80.001.00

For the Forger's own domains the dedicated adapter bdougie/smollm3-pokemon-forger-lora scores higher (body 0.82, outcome 0.66). Full numbers in eval.json; the training log in train.log.

Source code