MagicCaster/crimeradar-event-merge-qwen3-4b-reasoning-20260629
crimeradar-event-merge-qwen3-4b
A Qwen3-4B-Base student, full fine-tuned (no LoRA) to do event_merge for CrimeRadar: given a set of candidate events (ASR-derived dispatch/incident reports with id, summary and segment timelines), decide which ones refer to the same real-world incident and emit the merged grouping. Distilled from a dual-teacher trajectory set (GLM-5.2 + minimax) on the private dataset NewsBreak/crimeradar-event-merge (flat train split, both teachers). 4B counterpart of the qwen3-8b reasoning model.
Output format
The model is a reasoning model: it produces a <think>...</think> block, then a JSON answer with events_csv (and groups). The chat template places the opening <think> in the prompt, so generation begins inside the reasoning block and the model emits the closing </think> followed by the answer.
Eval (60-item consensus benchmark, vs 3-judge consensus gold)
Pairwise P/R/F1 + ARI, sampling temp 0.8 / topp 0.95 / topk 40 / presence_penalty 0.5.
A 4B student at teacher level on F1, ≫ the prior best 8B SFT (0.630). The larger qwen3-8b sibling reaches F1 0.824. Best single seed here: P 0.812 / R 0.886 / F1 0.848.
Recommended inference (vLLM)
Reliability-first sampling: temperature=0.8, top_p=0.95, top_k=40, presence_penalty=0.5, max_tokens=32768, stop on <|im_end|>. (These are baked into generation_config.json; presence_penalty is the one thing to pass at request time.)
A rare tail of the largest inputs (n_events≈11–15, with shared ASR segments / ambiguous boundaries) can run the <think> block to the token cap without closing it (~1% of generations). To bound it, serve with vLLM and a per-request thinking-token budget — the qwen3 reasoning parser already knows <think>/</think>:
vllm serve <this-repo> --reasoning-parser qwen3
# per request: SamplingParams(thinking_token_budget=16384) # offline
# extra_body={"thinking_token_budget": 16384} # OpenAI APIThis clips only the runaway loops and leaves normal reasoning untouched; F1 is unchanged.
Training
Full fine-tune, ZeRO-2, 8×GPU. 3 epochs (639 steps), lr 6e-5 cosine, warmup 0.05, effective batch 64 (per-device 1 × grad-accum 8 × 8 gpu), max_len 65536, bf16/tf32, flash-attention-2, gradient checkpointing, liger kernel. Final train loss ≈ 0.067. Trained on large-scene-enriched data (≥16-event inputs ~32% of training rows).
