wallfacers/weft-lineage-extractor-0.5b
weft-lineage-extractor-0.5b — smallest scale point of a NEGATIVE RESULT
## ⚠️ RESEARCH ARTIFACT — the 0.5B point of a synthetic-only training study. Not a production tool. ### ✅ Resolved by real data: use [weft-lineage-extractor-3b](https://huggingface.co/wallfacers/weft-lineage-extractor-3b) (real corpus, real precision 0.64). Full study: [weft-lineage-extractor-1.5b](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b).
The 0.5B point of a study showing that synthetic-only training induces a verbatim memorization leak in small models for ETL table-lineage extraction. It is the smallest scale point and shows the worst leak: near-perfect synthetic precision (0.994) collapses to 0.243 on real GitHub scripts, with 37.4% of hallucinations being table names recited verbatim from the synthetic training pool.
Same recipe as the 1.5B main model (LoRA on Qwen2.5-Coder-Instruct, Python/Shell synthetic ETL scripts, zero real scripts); only the base size differs.
This point's numbers (table-level, Convention A)
Where it fits (scale curve)
Leak shrinks with scale (capacity), but only real training data closes the real-world gap (bottom row). Direction confusion does not improve with size.
Intended use
- ✅ Reproducing / studying the synthetic-training memorization-leak failure at minimal scale.
- ❌ Not for production lineage — use the real-corpus 3B.
Usage, prompt format, training details, citation
Identical to the main model (this variant uses task_type: PYTHON | SHELL; everything else the same): [weft-lineage-extractor-1.5b](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b).
- Dataset & reports: wallfacers/weft-script-lineage-synth
- Base model: Qwen/Qwen2.5-Coder-0.5B-Instruct
- Platform: Weft (data-weave)
