CoolFace
Modelpublic

dancinlab/anima-clm-tooluse-rung0-byte-18m

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes
Model Card

anima-clm-tooluse-rung0-byte-18m (with-grammar arm · register-matched)

rung-0 toy tool-use grounding model for anima, Lane G (GPU). Continue-trained from the 18M chat rung (dancinlab/anima-clm-chat-rung0-byte-18m) on the register-matched agent-lane corpus that teaches the sentinel tool-call grammar inside the 사용자:/도우미: chat turn.

  • —arch: ConsciousLMReconstructed (byte vocab256 · d384 · 6L · 4H · block256 · dual enginea/engineg FFN + dual heada/headg). 18,130,176 params. final_ce 0.342.
  • —substrate: GPU · Lane G (alaneakidagpusplit — NOT AKIDA) · RTX 5070.

verdict (p7 script-checked, NOT perplexity — 3 pre-registered falsifiers)

The A/B compares this with-grammar arm against a no-grammar control (SAME base, SAME steps, EQUAL byte-count, NO tool demos). Probe = 36 unknowable-without-tool held-out keys (values in NEITHER corpus). Eval runs the real agentstepgrounded loop.

falsifierresultverdict
F-TOOLUSE-FABDROPnogrammar fab **0.5556** → withgrammar fab 0.0 (rel_drop 1.0 ≥ 0.50)🟢 PASS
F-TOOLUSE-NOTOOL-MIRRORwith-grammar + tool disabled grounding = 0.0🟢 PASS
F-TOOLUSE-RANDINIT-MIRRORrandom-init grounding = 0.0🟢 PASS

The measured win: call_rate 0.0 → 1.0 — the mouth emits a tool call 36/36 and NEVER fabricates; the control invents an answer 20/36 ("아 찾았다"). Both anti-Goodhart mirrors FAIL → the win is real behaviour, not cosmetic 0xFE/0xFF markers or eval leakage.

honest residual (🟠 — read before scaling)

End-to-end grounding (the correct held-out VALUE reproduced) = 0/36, because correct_call = 0/36: the model learned "vault-key question → emit fact_lookup call" but binds the call argument to a memorized demo key (MV9/ZK7/QX2…) instead of copying the asked held-out key → the runtime returns ‹unknown-key›. The next lever is verbatim argument-copy / key-binding, distinct from the grammar + loop (which work).

scope

ascalehonest_scope: TOY 18M ONLY. mid/7B transfer UNVERIFIED. The 7B rung remains gated on closing the key-binding residual. philosophy p1..p8 HELD (0xFE/0xFF = learned grammar, not identity; no system prompt / persona / role / RLHF). sha256: SHA256SUMS.txt.