CoolFace
Datasetpublic

wallfacers/weft-script-lineage-synth

weft-script-lineage — synthetic training data + negative-result evidence Companion data for the model wallfacers/weft-lineage-extractor-1.5b, a research artifact demonstrating that synthetic-only training induces a memorization leak in small models for ETL table-lineage extraction. Contents train.jsonl / heldout.jsonl — 10,000 + 600 synthetic ETL scripts with table-lineage labels (Python/Shell). out-jvm/ adds the Scala/Java-augmented variant. reports/ — the… See the full description on the dataset page: https://huggingface.co/datasets/wallfacers/weft-script-lineage-synth.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes38downloads
Dataset Card

weft-script-lineage — synthetic training data + negative-result evidence

Companion data for the model wallfacers/weft-lineage-extractor-1.5b, a research artifact demonstrating that synthetic-only training induces a memorization leak in small models for ETL table-lineage extraction.

Contents

  • —train.jsonl / heldout.jsonl — 10,000 + 600 synthetic ETL scripts with table-lineage labels (Python/Shell). out-jvm/ adds the Scala/Java-augmented variant.
  • —reports/ — the frozen evaluation evidence:
  • —eval-report-v2.md — synthetic held-out (near-perfect, the misleading number).
  • —eval-real.md / eval-real-jvm.md — four-way comparison on real GitHub scripts.
  • —leak-report.md / leak-report-jvm.md — memorization-leak quantification.
  • —leak-curve.md — leak vs. model scale (0.5B/1.5B/3B).
  • —paper-negative-result-findings.md — the full write-up.
  • —ADJUDICATION.md — the human-gold labeling convention (Convention A) + provenance.

Note on real scripts

The real-evaluation gold (realeval/gold/*.jsonl) contains third-party GitHub script content and is not included here by default (license/redistribution care). It is uploaded only when the publisher passes --include-real-gold after license review.