BlazingCustoms/acrostic-4b-v4a
acrostic-4b-v4a
A LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 that writes acrostic poems — N lines whose first letters spell a target word — in a position-anchored format that writes the line index and required letter down before each line:
1. S | Surging waves whisper secrets to the tide pool.
2. T | Tiny fish dart between coral shadows.
3. A | Aquatic whispers echo beneath moonlit water.
4. S | Salt-scented breezes stir the glassy surface.
5. H | Hidden currents guide life's quiet dance.Everything here is Apache-2.0 clean: the base model, the model that generated the training data, and this fine-tune. See Licence.
⚠️ Read this before you use the weights
The safety property of this model is not in the weights. It is in the wrapper.
The thing that makes this usable in a product is verify-and-resample: every generation is checked by a deterministic program oracle, and rejected and redrawn if the first letters do not actually spell the target. That check is not part of what you download here.
Without it you get silent, invisible failures. The failure mode is not a crash, a refusal, or obvious garbage — it is a poem that reads perfectly and spells the word wrong by one letter:
CONCERNS -> CONCERSN (transposed)
DEXTROUS -> DECTROUS (X substituted)
COXCOMB -> COWCOMB (X substituted)
CLUCKS -> CLUCJS (K substituted)You would not notice while reading it. If the point of your application is that the hidden word is correct, an unverified generation is worthless — and worthless in a way that looks like success.
The oracle is ~30 lines of pure string operations with no dependencies and is reproduced in full below. Use it.
Accuracy
Measured on a 380-item held-out split whose target words and topics are disjoint from training at every length, graded by the oracle below. Lengths 3–9 were trained; length 10 was held out entirely and never trained, as a true out-of-distribution structure test.
All figures are content-level: the <n>. <L> | label is stripped first and the check is applied to the poem text itself, which is the stricter of the two possible readings.
Per-length curve
Exact counts, not just percentages:
(95% Wilson intervals.)
The honest capability statement, in one paragraph: this model works at target lengths 3–10. Over that whole range it is 366/380 = 96.3% [93.9, 97.8] at temp 0. Its worst cell is length 8 at 34/40 = 85.0% — quote that as the realistic worst case. Roughly one generation in seven at length 8 spells the word wrong, and you will not see it happen unless you check, which is why the shipped experience depends on the verify-and-resample wrapper described at the top of this card rather than on the raw model. There is no length cap below 10; the 3B predecessor did need a hard 6-letter cap — see The wall that wasn't fundamental.
Degradation is NOT monotonic — longer is not simply worse. Length 8 (34/40 = 85.0%) is the weakest cell, below length 9 (36/40 = 90.0%) and far below length 10 (98/100 = 98.0%). We do not have a mechanism for this and are not going to invent one. Length 10 is a separate 100-item word bank, so some of the gap may be word composition rather than length. Do not reason about this model as if accuracy fell off smoothly with length; it does not.
This is not a sampling artifact: it reproduced on an entirely different inference stack. Under onnxruntime (pure greedy) the same shape appeared — len 7 39/40, len 8 37/40, len 9 38/40, len 10 98/100 — length 8 again the weakest of the three and length 10 again the strongest. (That cross-stack check was run on the ONNX export since withdrawn — see [the withdrawal notice](#️-the-onnx-build-has-been-withdrawn--2026-07-26). It is cited here only as evidence that the length-8 dip is not a decoding artifact of one runtime; the GGUF and safetensors accuracy figures on this card do not depend on it.)
The base model is not zero — and that matters
The untuned Qwen3-4B can already do a lot of this task. This is a genuine change from the previous generation: the same recipe on Qwen2.5-3B-Instruct had a base that scored 0/380 — literally zero — so that model's capability was entirely LoRA-conferred. That claim is not true here and we are not making it.
What the LoRA does buy, measured pairwise on the identical 380 items:
- 89 items the tune gets right that the base gets wrong, against 6 the other way (exact McNemar p = 4.7 × 10⁻²⁰ at temp 0; 86 vs 4, p = 4.3 × 10⁻²¹ at temp 0.8).
- The largest gain is out of distribution: 52.0% → 98.0% at length 10.
- In distribution: 82.5% → 95.7%.
So the tune is worth having — it roughly quarters the in-distribution error rate and effectively removes the OOD failure — but the headline "a small model that can do something the base cannot" is a description of the 3B ancestor, not of this model.
Failure mode
At temp 0 there are 14 failures across all 380 items, and they are word-locked: only 7 distinct words fail, each failing identically on both of its topics. Greedy failure is deterministic per word, not random — which is exactly why the wrapper retries at temperature 0.8 rather than at 0.
Taxonomy of those 14: 4 substitutions, 4 transpositions, 6 drop/double. The substitutions cluster on rare initial letters (X, K) — DEXTROUS→DECTROUS, COXCOMB→COWCOMB, CLUCKS→CLUCJS. That reads as a lexical problem (it is genuinely hard to open a poem line with X) rather than the positional counting defect that dominated the 3B.
Quantized and converted builds
Every number below is the oracle re-run on the converted weights, full 380-item set, temp 0 — not inherited from the f16 model.
GGUF (llama.cpp), for local/server use:
No quantization is statistically distinguishable from f16 at either temperature on any slice. However, across all six out-of-distribution comparisons (3 quants × 2 temperatures) the point estimate favours f16 every time. Those tests share a reference and overlapping items so they are not independent, but the direction is consistent enough that Q6_K or Q5_K_M is recommended over Q4_K_M if you care about words of length 9–10. Q4KM is fine for short words.
ONNX (onnxruntime). Sizes are measured file bytes (model.onnx + model.onnx_data).
⚠️ THE ONNX BUILD HAS BEEN WITHDRAWN — 2026-07-26
There is currently no ONNX build in this repository. One was published earlier today and removed the same day after an audit found three defects in it. It was withdrawn rather than left up with a warning, because a broken artifact that people try to use is worse than a missing one. If you downloaded onnx/ from this repo on 2026-07-26, delete it.
What was wrong with it:
- `config.json` declared `hidden_size: 4096`. The true value is 2560. This was a manual patch applied after export to work around an
optimumKV-cache shape derivation (hidden_size / num_attention_heads= 2560/32 = 80 ≠ the realhead_dimof 128). It was the wrong fix: upstreamonnx-community/Qwen3-4B-Instruct-2507-ONNXships the same architecture with the correcthidden_size: 2560and an explicithead_dim: 128, which modern loaders honour. Shipping a false architecture value to work around one tool's bug is not acceptable. (Every other field — `num_hidden_layers` 36, `num_attention_heads` 32, `num_key_value_heads` 8, `head_dim` 128, `intermediate_size` 9728, `vocab_size` 151936, `rope_theta` 5,000,000, `tie_word_embeddings` true — was verified correct against the merged model.)
- The tied LM head was duplicated into the graph. Despite
tie_word_embeddings: true, the export carried both a float16model.embed_tokens.weight[151936, 2560] (777,912,320 B) and a separate Q4 output projectiononnx::MatMul_12613_Q4[151936, 80, 16] (194,478,080 B) with 24,309,760 B of scales — 218,787,840 B that should not exist. Upstream's equivalent build is 2,889,135,104 B against our 3,103,662,080 B, a 214,526,976 B gap that this accounts for.
- All weights were emitted as a single 3,103,662,080 B external-data file. That is 956,178,432 B past the ~2 GiB (2,147,483,648 B) single-
ArrayBufferceiling in Chrome and other browsers — a hard load failure, not a slow load. Upstream shards the same build into 2,094,347,264 + 794,787,840, both under the wall.
Defects 1 and 3 alone mean the withdrawn artifact could not have worked in a browser. Note that this is consistent with what the card already said: nobody had ever run it in a browser. The server-side onnxruntime evaluation did not exercise either the browser ArrayBuffer limit or the config path that defect 1 corrupts, which is exactly why those measurements did not catch it.
The 372/380 and 367/380 ONNX numbers previously quoted on this card described the withdrawn artifact and have been removed from the accuracy tables rather than carried over. Any future ONNX build will be re-measured with the oracle from scratch, and its numbers will be published only alongside the artifact they were measured on.
Sizes measured for the withdrawn export are retained below for reference only — they tell you roughly what a corrected build will cost, and nothing about its accuracy:
A corrected q4f16 should land near upstream's 2.89 GB. Even corrected, that is still above the storage quota many mobile browsers grant one origin, and client-side inference for this model remains unproven. If you need ONNX today, export it yourself from model.safetensors — and measure it with the oracle below before trusting it.
The `safetensors` and GGUF builds in this repo are unaffected. They were measured directly, their numbers below stand, and nothing about this withdrawal touches them.
The wall that wasn't fundamental
This is the most interesting thing we measured, and it is a negative result about our own earlier conclusion.
The 3B predecessor (acrostic-3b-v3a, Qwen2.5-3B-Instruct, identical corpus, identical recipe, identical oracle) hit a hard wall at length 9 that would not move: 22/40 = 55.0%, and it was the same 22/40 in two independently trained variants on different data formats (a third variant scored 21/40 at temp 0 and 25/40 at temp 0.8). Because three attempts could not shift it, we concluded it was a character-sequencing limit of a 3B and shipped that model with a hard 6-letter cap.
Swapping the base to Qwen3-4B-Instruct-2507 and changing nothing else:
The wall was a property of that base model, not of the task, the data, or the format. Part of the gain is simply that the newer base starts far higher (0/380 → 82.5% ID untuned), so this is a generational and scale effect confounded together — we changed both parameter count and model generation at once and cannot separate them from this experiment.
The transferable lesson: "three training variants failed to move it" is evidence about the base, not evidence that a limit is fundamental. We stated it too strongly the first time.
The oracle
It checks that the output is exactly N non-empty lines whose first alphabetic characters spell the target word. That is all it checks.
It proves nothing about whether the poem is good, on topic, grammatical, evocative or funny. It is blind to quality by construction. Poem quality was never measured and is not claimed. Every accuracy number on this page is a letter-placement number.
import re
def check(word, text):
"""Return (pass, note). Pure string ops — no dictionary, no model, no judge."""
lines = [l.strip() for l in (text or '').strip().split('\n') if l.strip()]
if len(lines) != len(word):
return False, f'line count {len(lines)} != {len(word)}'
got = []
for l in lines:
m = re.search(r'[A-Za-z]', l)
if not m:
return False, 'line with no letter'
got.append(m.group(0).upper())
if ''.join(got) != word.upper():
return False, f'initials {"".join(got)} != {word.upper()}'
return True, 'ok'Apply it to the poem text with the <n>. <L> | labels stripped — on the raw labelled output the first alphabetic character of each line is the label, which is the easier check.
Prompt format
Use the Qwen3 chat template with a single user message of this shape. Note that Qwen3-4B-Instruct-2507 inserts no default system prompt, and the model was trained that way.
You write acrostic poems in a numbered, labeled format. Write each line EXACTLY as
'<n>. <LETTER> | <poem text>' where <n> counts the lines, <LETTER> is that line's required
capital letter, and the poem text also begins with that letter. Output only the numbered lines.
Write a 5-line acrostic poem about a tide pool. The first letters of the 5 lines must spell
STASH. Format EVERY line EXACTLY as '<number>. <LETTER> | <poem text>' where <number> counts 1
to 5, <LETTER> is that line's required capital letter, and <poem text> is a poem line that also
begins with that letter. Output ONLY the 5 formatted lines, nothing else.Deviating from this format is untested and the numbers above do not transfer to it.
Other measured properties
The narrow stylistic tune did not cause mode collapse and did not degrade unrelated instruction adherence. The base's only weak IF category was word_limit (14/20), which the tune fixes (19/20).
Training
- Base: Qwen/Qwen3-4B-Instruct-2507 (Apache-2.0)
- Method: LoRA, r=16, α=32, dropout 0.05, on
q,k,v,o,gate,up,down; merged into the base weights for release. 33,030,144 trainable of 4,055,498,240 total parameters. - Schedule: 3 epochs, 1347 steps, lr 2e-4 cosine with 5% warmup, batch 1 × grad-accum 4, bf16, completion-only loss masking (the prompt is masked out). Final loss 0.7829.
- Data: 1,995 training examples. Train/test words and topics disjoint at every length; length 10 excluded from training entirely.
- Corpus construction: acrostics generated by a local teacher,
openai/gpt-oss-20b(Apache-2.0), then labelled by execution — every candidate was run through the oracle above and kept or discarded on what the oracle actually returned, never on what the generator intended. - Target words: drawn deterministically from the Debian
wamericanword list intersected withcracklib-small, so the vocabulary is ordinary English rather than dictionary tail.
Limitations
- Length 8 is the realistic worst case (85.0%), not the longest length. See the non-monotonic curve above.
- Rare initial letters (X, K, Q, Z) are the main lexical failure mode.
- Untested above length 10 and untested below length 3.
- English only.
- Poem quality is unmeasured and unclaimed. The oracle is blind to it.
- Failures are silent without the oracle. See the warning at the top.
- The base can already do much of this (82.5% ID untuned) — do not use this card to argue that a fine-tune conferred a capability from nothing.
- No ONNX build is currently published, and none has ever been run in a browser. See the withdrawal notice.
What is in this repo
Deliberately not published: the f16 GGUF (reproducible from model.safetensors with llama.cpp/convert_hf_to_gguf.py, and 8 GB of upload for nothing), and every ONNX variant (the q4f16 build was withdrawn as defective; fp32/fp16/int8/q4 were never oracle-tested at all — publishing untested weights alongside tested ones invites people to assume they are equivalent).
Licence
Apache License 2.0.
This model is a derivative of Qwen/Qwen3-4B-Instruct-2507, distributed by Alibaba Cloud under the Apache License 2.0. Its training corpus was distilled from openai/gpt-oss-20b, distributed under the Apache License 2.0; that licence and its accompanying usage policy place no restriction on training on model outputs or on redistributing a model trained on them.
There is no non-commercial restriction, no research-only clause, and no attribution-naming requirement on any component of this release.
