CoolFace
Modelpublic

elifepathways/jats-agentic-annotation-qwen3.5-9b-teacher-distill-v1

sourceHugging Faceapache-2.0updated 26d agoView on Hugging Face
0likes53downloads
Model Card

JATS agentic annotation — Qwen3.5-9B teacher-distilled LoRA (v1, "tfull")

Full merged weights (Qwen3.5-9B + teacher-distilled LoRA) for a multi-turn tool-call agent emitting JATS XML for scientific-paper body sections via the scratchpad environment in agentic-jats-annotation-inference.

These are ready-to-serve full weights (LoRA already merged onto the v11 SFT base). The raw adapter is also included under lora_adapter/ — if you use it instead, apply it on top of the v11 SFT merge, NOT raw Qwen3.5-9B (measured: 0.234 vs 0.452 pass@6 body F).

Training

  • —Teacher: DeepSeek-v4-pro driven turn-by-turn through the 17-tool JatsEnv on sup/sub-preserving markdown (docling + custom run/footnote/ins fixes); 1,549 rollouts over 800 docs (~$250 API), filtered to best-per-doc with F0.5 ≥ 0.45 and recall ≥ 0.30 (~430 trajectories).
  • —SFT: LoRA r32/α64, lr 2e-5, 2 epochs, max-length 16384, assistant-token masking, bf16, sdpa attention (flashattn2 crashes on this architecture).

Evaluation (77-doc fixed body slice, strict exact-match metric, pass@6 @ T=0.9)

modelplain F+ candidate mergemedian Fmean single-rollout F
v11 SFT base0.3360.5110.2580.172
+ 215-doc teacher SFT0.3720.5360.3400.148
this adapter (~430 docs)0.4520.5950.4880.180

Paired vs base: +0.118 (95% CI +0.067..+0.169). Perfect-play ceiling under this metric is 0.842. "Candidate merge" = deterministic DOCX-style-run and regex-citation candidates union-merged post-hoc at serialize time (scripts/candidate_merge.py in the training repo); recommended in production. PDF-sourced inputs score lower (≈0.16 plain).

Serving

vLLM ≥ 0.19, --max-model-len 41472, thinking mode ON, temp 0.7–0.9; stop on </tool_call>, one call per turn, env owns the XML tree. See the inference repo for the full annotation loop. Eval evidence: eval_passk_tfull_v2slice.json (per-doc f_lists + full traces).

Per-tag evidence (also in this repo)

  • —eval_pertag_comparison.json — base/tpwarm/tfull with paired bootstrap CIs (tfull vs base: +0.118 [+0.066,+0.172] plain; docx 0.487/0.632, pdf 0.294/0.429).
  • —eval_pertag_by_tag.json — per-JATS-tag P/R/F1 of best rollouts, with and without candidate merge. Highlights (F1, plain → merged): sub 0.00 → 0.95, sup 0.34 → 0.54, xref 0.33 → 0.65.

Note for LoRA-only users: adapter_config.json:base_model_name_or_path refers to a local training path; the required base is the v11 scratchpad-SFT merge of Qwen/Qwen3.5-9B described above (not distributed here — contact the authors, or use the merged weights release when available).