aimeri/spoomplesmaxx-thrasher-24B
<!doctype html> <html lang="en"> <head> <meta charset="UTF-8" /> <meta name="viewport" content="width=device-width, initial-scale=1.0" /> <title>SpoomplesMaxx Thrasher 24B</title> </head> <style> @import url("https://fonts.googleapis.com/css2?family=Consolas&display=swap"); .crt-container { padding: 10px; max-width: 1000px; margin: 0 auto; width: 95%; } .crt-case { background: #e8d7c3; border-radius: 10px; padding: 15px; box-shadow: inset -2px -2px 5px rgba(0, 0, 0, 0.3), 2px 2px 5px rgba(0, 0, 0, 0.2); } .crt-inner-case { background: #e8d7c3; border-radius: 8px; padding: 3px; box-shadow: inset -1px -1px 4px rgba(0, 0, 0, 0.3), 1px 1px 4px rgba(0, 0, 0, 0.2); } .crt-bezel { background: linear-gradient(145deg, #1a1a1a, #2a2a2a); padding: 15px; border-radius: 5px; border: 3px solid #0a0a0a; position: relative; box-shadow: inset 0 0 20px rgba(0, 0, 0, 0.5), inset 0 0 4px rgba(0, 0, 0, 0.4), inset 2px 2px 4px rgba(255, 255, 255, 0.05), inset -2px -2px 4px rgba(0, 0, 0, 0.8), 0 0 2px rgba(0, 0, 0, 0.6), -1px -1px 4px rgba(255, 255, 255, 0.1), 1px 1px 4px rgba(0, 0, 0, 0.3); } .crt-bezel::before { content: ""; position: absolute; top: 0; left: 0; right: 0; bottom: 0; background: linear-gradient( 45deg, rgba(255, 255, 255, 0.03) 0%, rgba(255, 255, 255, 0) 40%, rgba(0, 0, 0, 0.1) 60%, rgba(0, 0, 0, 0.2) 100% ); border-radius: 3px; pointer-events: none; } .terminal-screen { background: #140d06; padding: 20px; border-radius: 15px; position: relative; overflow: hidden; font-family: "Consolas", monospace; font-size: clamp(12px, 1.5vw, 16px); color: #ff9e3d; line-height: 1.4; text-shadow: 0 0 2px #ff9e3d; filter: brightness(1.1) contrast(1.1); box-shadow: inset 0 0 30px rgba(0, 0, 0, 0.9), inset 0 0 8px rgba(0, 0, 0, 0.8), 0 0 5px rgba(0, 0, 0, 0.6); max-width: 80ch; margin: 0 auto; } .terminal-screen h2, .terminal-screen h3 { font-size: clamp(16px, 2vw, 20px); margin-bottom: 1em; color: #ffd23f; text-shadow: 0 0 3px rgba(255, 210, 63, 0.5); } .terminal-screen pre.code-block-image { display: inline-block; text-align: left; font-size: clamp(2px, 0.4vw, 12px); font-family: monospace; margin: 1em 0; background-color: #1a1a1a; padding: 1em; border-radius: 4px; color: #ff9e3d; overflow-x: auto; line-height: 1; max-width: 100%; white-space: pre; } .terminal-screen pre.code-block { display: inline-block; text-align: left; font-size: clamp(10px, 1.3vw, 14px); font-family: monospace; margin: 1em 0; background-color: #1a1a1a; padding: 1em; border-radius: 4px; color: #ff9e3d; overflow-x: auto; line-height: 1; max-width: 100%; white-space: pre; } .terminal-screen::before { content: ""; position: absolute; top: 0; left: 0; right: 0; bottom: 0; background: linear-gradient( rgba(18, 16, 16, 0) 50%, rgba(0, 0, 0, 0.25) 50% ), url("data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAADIAAAAyBAMAAADsEZWCAAAAGFBMVEUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA4o8JoAAAAB3RSTlMAGwQIEQMYADcPzwAAACJJREFUKM9jYBgFo2AU0Beg+A8YMCLxGYZCbNQEo4BaAAD5TQiR5wU9vAAAAABJRU5ErkJggg=="); background-size: 100% 2.5px; pointer-events: none; z-index: 2; } .terminal-screen::after { content: ""; position: absolute; top: 0; left: 0; right: 0; bottom: 0; background: radial-gradient( circle at center, rgba(20, 13, 6, 0) 0%, rgba(20, 13, 6, 0.2) 50%, rgba(20, 13, 6, 0.15) 100% ); border-radius: 20px; pointer-events: none; z-index: 1; } .terminal-screen .notice { margin: 1.5em 0; padding: 0.8em 1.2em; border: 1px solid #ffd23f; border-radius: 4px; background-color: rgba(255, 210, 63, 0.04); } .terminal-screen .notice h3 { margin-top: 0.2em; margin-bottom: 0.5em; } .terminal-screen .notice p { margin-bottom: 0.2em; } .terminal-screen strong, .terminal-screen em { color: #f0f0f0; } .terminal-screen p, .terminal-screen li { color: #ff9e3d; } .terminal-screen a { color: #6fb3ff; text-decoration: underline; text-shadow: 0 0 2px rgba(93, 169, 255, 0.5); transition: opacity 0.2s; } .terminal-screen a:hover { opacity: 0.8; } .terminal-screen code, .terminal-screen kbd, .terminal-screen samp { color: #ff9e3d; font-family: "Consolas", monospace; text-shadow: 0 0 2px #ff9e3d; background-color: #1a1a1a; padding: 0.2em 0.4em; border-radius: 4px; } </style> <div class="crt-container"> <div class="crt-case"> <div class="crt-inner-case"> <div class="crt-bezel"> <div class="terminal-screen"> <div style="text-align: center"> <h2>SpoomplesMaxx-Thrasher-24B</h2> <h3>"Thrash Metal"</h3> <pre class="code-block-image"> ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ░░░░░ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░██ ▒▓██▓▒ ░ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░ ▓░▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░▓▓▒█▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ░░░░░█▓░▓█ ░▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░▒███▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▒░░▒░░░░▓ ▓ ▓█▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░░░░▓░ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░░░▒░░ █▓ █▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ░░░░░░░▒░░▒▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░▒░░░░ █ ▓▓▓ ▓▓█▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░ ░░░░░ ▒░▓▓░ █▓▓▒▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░ ▓▓ ▓ ▓ ▓▓▓ ▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ █▓▓ █▓▓▓ ▓▓█▓█ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▒ █ ▓ ░▓▓▓▓ ▓▒▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▒ ▓ ▓▒▓▓ ▒▓▓▓▓▓▓▓█ ▓▓▒▓▓▓▓▓▓▓▓▓▓▓▓▓░ ▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░ ▒ ▒ ▓ ▓▓ ▓▓ ▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ░ ░ ░░▓▓▓▓ ▓▓░▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ░ ░░░▓▒█▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓░░░░▓▓ ▓▓▓▓ ░▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░▓ ▓▓▓░ ▓▓▓▓ ▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ░▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓█▓▓▓▓▓ ▓▓▓▓▓▓▓▓ ▒▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓ ░▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▒▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓░ ░ ▒▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓ ░ ▓ ▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓ ▒▓ ▒ ▓▓▓▓▓▓▓▓▓▓▓ ░▓▓▓▓▓▓▓ ▓ ▒▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓ ▓ ░▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓ ▒ ▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓ ▓ ░▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓ ░▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓░▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓░ ▓▓ ▓▓ ▓ ░░▓▓▓▓▓▓▓▓▓▓▓ ▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓ ▓ ░ ▓▓▓▓▓▓▓▓▓ ▒▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓ ░░▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓ ▒ ░ ▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓ ░ ░▓ ▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓ ░▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓░ ▓ ▒▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ </pre> </div> <p><strong>"Thousand-Songed"</strong> — <em>Toxostoma rufum</em>, the brown thrasher: holder of the largest documented song repertoire of any North American bird, over a thousand song types. The mockingbird repeats a phrase three times; the thrasher sings each one twice and moves on. Second of the mimids, the family that follows the corvids.</p> <p><strong>A model that is 100% about roleplay, trained mostly on things that are not roleplay.</strong> Measured across the strongest open RP lineage I know (<a href="https://huggingface.co/PocketDoc/Dans-PersonalityEngine-V1.3.0-24b">Dans-PersonalityEngine</a>), roughly 590K of its rows are task/reasoning/assistant/world-knowledge data against ~150K of actual roleplay. RP is the product; RP is not the corpus. The RP data teaches the register. Everything else teaches the mind behind it.</p> <p>Built on <strong>Mistral-Small-3.1-24B-Base</strong> with the vision tower removed and the chat interface rebuilt from scratch — see "The token surgery" below, because if you have ever bounced off a Mistral model's template, that section is for you.</p> <h3>Who this is for</h3> <p>The accessible mimid. mockingbird at 36B asks a lot of your VRAM; thrasher at 24B (23.6B after the vision strip) is the same recipe in a size that quantizes onto a single 24GB card. If you can run any of the popular 22–24B RP models, you can run this one.</p> <h3>Prompt format</h3> <p><strong>ChatML.</strong> On a Mistral base. Yes, really — and not by resizing anything:</p> <pre class="code-block"> FORM <|imstart|>system\n{card}<|imend|>\n<|imstart|>user\n{text}<|imend|>\n<|imstart|>assistant\n{reply}<|imend|>\n STOPS <|imend|> (id 21; eos) BOS <s> (id 1) — added automatically by the tokenizer, not the template ROLES system / user / assistant / tool EXAMPLE <|imstart|>system You are Bram Hollis, keeper of the Wayward Lantern...<|imend|> <|imstart|>user I push the door open, dripping wet Got room for one more?<|imend|> <|imstart|>assistant </pre> <p>The template ships embedded (<code>chattemplate.jinja</code> + <code>tokenizerconfig.json</code>), so vLLM, llama.cpp, MLX, and every frontend that speaks ChatML — which is to say, every RP frontend — picks it up without ceremony.</p> <div class="notice"> <h3>No thinking. Ever.</h3> <p>thrasher never emits reasoning traces and was never trained on them.</p> </div> <h3>The token surgery</h3> <p>Moving a Mistral base to ChatML was my first barrier the first time I ever tried to use a Mistral model for fine-tuning, so here is the whole recipe, with receipts.</p> <p><strong>1. There is room in the vocabulary.</strong> Mistral's Tekken tokenizer reserves ids 0–999 as a control block; only 0–19 are named. Ids 20–999 are unused <code><SPECIALn></code> placeholders. So ChatML needs <strong>no vocab resize and no embedding-matrix growth</strong>: rename <code><SPECIAL20></code> → <code><|imstart|></code> and <code><SPECIAL21></code> → <code><|imend|></code> in <code>tokenizer.json</code> (both <code>addedtokens</code> and the vocab), <code>tokenizerconfig.json</code>, and <code>specialtokensmap.json</code>. Set eos to <code><|imend|></code>. Done — single-token ids 20 and 21.</p> <p><strong>2. The claimed rows are dead, and dead eos rows are fatal.</strong> The base was pretrained with those placeholders never appearing in data: their embedding rows are <strong>exactly 0.0</strong> (measured), their lmhead rows random-scale noise. A model whose eos row is dead never learns to stop — unless you train very long (Hermes cold-claimed on 60B tokens; this SFT is ~1B) or you initialize sensibly. thrasher grafts: <code><|imstart|></code> rows ← <code><s></code>, <code><|imend|></code> rows ← <code></s></code>, embedding <strong>and</strong> lmhead. By step 100 of SFT, stop-rate at temperature 0.7 was already 24/24; it stayed 100% at every checkpoint probed.</p> <p><strong>3. The published tokenizer has a broken pre-tokenizer regex.</strong> The known Mistral conversion bug (transformers warns and offers <code>fixmistralregex=True</code> — but that fix is in-memory only, and axolotl, llama.cpp, and MLX read <code>tokenizer.json</code> directly). The true Tekken pattern — case-aware word splitting, single-digit number splits — is baked into the shipped file.</p> <p><strong>4. The vision tower is gone.</strong> 222 tensors of Pixtral removed, the <code>languagemodel.</code> prefix stripped, published as a plain untied <code>MistralForCausalLM</code>. Nothing multimodal remains.</p> <p>The full prep script (<code>prepbase.py</code>) is in this repo, and the prepped base is its own artifact if you want to start from it: <a href="https://huggingface.co/aimeri/Mistral-Small-3.1-24B-Base-thrasher">Mistral-Small-3.1-24B-Base-thrasher</a>.</p> <h3>Tool calling</h3> <p>The corpus includes the full Toolmaxx family (58,095 conversations), rendered with tool responses as a plain <code>tool</code> role turn:</p> <pre class="code-block"> <|imstart|>tool\n{tool output}<|imend|> </pre> <div class="notice"> <h3>Corpus-taught, conversational tool competence — not a structured calling API.</h3> <p>If you need strict function calling, put a schema in the card and validate what comes back.</p> </div> <h3>Key details</h3> <pre class="code-block"> BASE mistralai/Mistral-Small-3.1-24B-Base-2503 (Apache 2.0, vision stripped) PARAMS 23.6B dense · 40 layers · GQA 8 KV heads · headdim 128 · hidden 5120 VOCAB 131,072 · ChatML on claimed Tekken slots 20/21 · zero added tokens CTX trained at 24,576 packed · base RoPE (theta 1e9) to 131K CORPUS 667,332 conversations · ~1.3B supervised chars · 43% RP share LANGUAGE English (non-English filtered at ingest; base priors remain) </pre> <h3>Training</h3> <p>Full-parameter SFT, <a href="https://github.com/axolotl-ai-cloud/axolotl">Axolotl</a>, 8×H200. One stage. The exact config generator ships in this repo (<code>thrashersftcfg.py</code>):</p> <pre class="code-block"> STEPS 902 (2 epochs) · this release = step 600 SEQ 24,576 · sample packing (block-masked; packing does not shrink context) BATCH 64 global (micro 1 × accum 8 × 8 GPUs) OPT AdamW · lr 8e-6 cosine · 3% warmup · wd 0.01 · bf16 STACK FSDP2 full-shard · activation checkpointing · Cut Cross Entropy HEALTH gradnorm 1.2–1.8 the whole run · zero spikes · memory flat </pre> <p><strong>On context:</strong> 24,576 is the longest single training conversation (longer ones were split at turn boundaries with the card re-carried). The base's 131K RoPE survives SFT untouched; the best-trained RP region is the first ~24K, degrading gracefully beyond.</p> <p>The corpus is mockingbird's, verbatim — the PersonalityEngine V1.3.0 public list plus my own carded-RP, think-stripped-RP, and anti-repetition lanes, same cleaning receipts. Only the template changed. See the <a href="https://huggingface.co/aimeri/spoomplesmaxx-mockingbird-36B">mockingbird card</a> for the full corpus story.</p> <h3>How the checkpoint was chosen — and how you can check</h3> <p>Loss did not pick this model. Checkpoints went through two instruments, <strong>both published in this repo's <code>eval/</code></strong>: a seeded multi-turn loop/stall battery (six gates, six repeats per episode — single-run numbers on it are noise, and the trend proves it), and blind-judged episodes on four real character cards.</p> <pre class="code-block"> step 100 200 300 450 600 750 900 battery 17 8 13 13 18 10 16 (of 24) stop-rate 100% 100% 100% 100% 100% 100% 100% </pre> <p style="text-align:center"><img src="https://huggingface.co/aimeri/spoomplesmaxx-thrasher-24B/resolve/main/eval/curves/lossvsbattery.png" alt="train loss keeps falling while battery pass rate peaks at step 600 and regresses" style="max-width:100%; border-radius:4px;" /></p> <p>The shape is the story: a mid-run dip while style reorganizes under high LR, a peak mid-way through epoch 2, then regression in the deep anneal. Behavioral peak ≠ end of training. <strong>Step 600 shipped</strong> — highest battery pass rate and the most disciplined judged transcripts (compressed, declarative, no user-impersonation).</p> <p><strong>The anti-repetition anneal experiments — published, not spun.</strong> After training, we annealed checkpoints 600 and 900 on a <a href="https://huggingface.co/datasets/aimeri/repremover-xl">7,281-conversation anti-repetition dataset</a> (recipe in that repo), across four blend/base variants plus two task-arithmetic merges. In our measurements, one variant reached zero repeat failures on the battery while regressing on format gates; the merges regressed stop-token reliability; none beat the un-annealed step 600 overall, so step 600 shipped as-is. The raw battery JSONs for every variant are in <code>eval/</code> — read them and draw your own conclusions rather than taking ours.</p> <h3>Sampling</h3> <p>The shipped <code>generationconfig.json</code> is the swept optimum (7 arms × 5 seeded battery runs each):</p> <pre class="code-block"> temperature 1.0 · minp 0.05 · topp off </pre> <p>thrasher is the anti-mockingbird in its sampler behavior, which is why we sweep per model instead of inheriting: minp won here (mockingbird's sweep found minp looped MORE); and a mild <code>repetitionpenalty 1.05</code> — catastrophic on mockingbird — is a legitimate opt-in on thrasher: in our sweep it eliminated the verbatim-loop tail entirely (max cross-turn Jaccard 0.50 across 20 episodes) at the cost of a rare unfinished turn. If loops bother you more than an occasional run-on, add it. Plain topp 0.9 at temp 1.0 was the <em>worst</em> repetition arm on this model — don't ship RP muscle memory, sweep.</p> <div class="notice"> <h3>Known limitation: the loop tail.</h3> <p>Under sustained low-information multi-turn pressure the model can fall into near-verbatim self-repetition — roughly 1–2 episodes in 24 on our battery at shipped settings. The corpus's anti-repetition lane suppresses it; it is not eliminated. The repetitionpenalty 1.05 opt-in above removed it entirely in our measurements.</p> </div> <p>Give it a proper card and it will give you a proper character: the model was fed real character cards (median ~3K chars, p90 ~8.5K) as system messages.</p> <h3>Quickstart</h3> <pre class="code-block"> from transformers import AutoModelForCausalLM, AutoTokenizer modelid = "aimeri/spoomplesmaxx-thrasher-24B" tok = AutoTokenizer.frompretrained(modelid) model = AutoModelForCausalLM.frompretrained(modelid, torchdtype="bfloat16", devicemap="auto") messages = [ {"role": "system", "content": "You are Bram Hollis, keeper of the Wayward Lantern... Third person, *asterisk action beats*."}, {"role": "user", "content": "*I push the door open, dripping wet* Got room for one more tonight?"}, ] ids = tok.applychattemplate(messages, addgenerationprompt=True, returntensors="pt").to(model.device) out = model.generate(ids, maxnewtokens=400, dosample=True) # sampler ships in generationconfig print(tok.decode(out[0][ids.shape[-1]:], skipspecialtokens=True)) </pre> <p>Quants: <a href="https://huggingface.co/aimeri/spoomplesmaxx-thrasher-24B-GGUF">GGUF static</a> (Q3/Q4/Q5KM) · <a href="https://huggingface.co/aimeri/spoomplesmaxx-thrasher-24B-i1-GGUF">GGUF imatrix</a> (IQ3XXS–Q4K_M, own-corpus calibration, imatrix.dat included) · MLX <a href="https://huggingface.co/aimeri/spoomplesmaxx-thrasher-24B-mlx-4bit">4-bit</a> / <a href="https://huggingface.co/aimeri/spoomplesmaxx-thrasher-24B-mlx-6bit">6-bit</a>. Prefer the imatrix quants at 3–4 bit.</p> <p>The thrasher knows a thousand songs. It only needs the one you hand it.</p> <p><em>thrasher is a roleplay and creative-writing model for adults. It stays in character by design — its corpus was scrubbed of mid-scene refusals — so bring your own moderation where your deployment needs it. Not an assistant, not an oracle, not for anything safety-critical.</em></p> <p><em>mimids 02 · trained 2026-08 · checkpoints at <a href="https://huggingface.co/aimeri/thrasher-v1-ckpts">thrasher-v1-ckpts</a> · eval instruments, prep scripts, and training config in this repo · Apache 2.0</em></p> </div> </div> </div> </div> </div> </html>
