Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-v2-h4-320
episodic-ingestion-compiler / ModernBERT field-event ranker (mixed-mode v2, H4-320)
Fine-tune of answerdotai/ModernBERT-base for the grouped multi-positive softmax ranker task, trained on v2 semantic-reasoning labels for 320 optimizer steps (vs 160 in the initial v2 checkpoint). The extra training let the model fully converge on both legacy and semantic fields.
Versus v2 at 160 steps
Training details
- Train rows: 1217 (mixed-mode v2)
- Train groups: 5882
- Eval rows: 220 (21986 candidate pairs)
- Steps: 320 optimizer steps × accum=8 = 2560 forwards
- Optimizer: AdamW lr 5.7e-5, warmup 30
- Precision: bf16 autocast + gradient checkpointing
- Peak VRAM: 4.08 GiB
- Hardware: RTX 5090 Laptop (Blackwell SM 12.0, 24 GiB)
Eval metrics
Per-field MRR (top fields in eval):
Lineage
Semantic-reasoning fields show clear learning
The fields designed to NOT be role-tautological reached high MRR through training, confirming the model is actually reasoning beyond role:
initiating_command: MRR 0.966 (3.51x random) — learns trace boundariesoutcome_of_latest_attempt: MRR 0.966 (3.50x random) — matches call_id chainsrecent_error: MRR 0.826 (3.23x random) — latest-failure reasoningaction_causality: MRR 0.561 (2.21x random) — content grounding across turns
Legacy fields (attempted_actions, observed_outcomes) sit near 0.48, close to random for multi-positive softmax. They're role-tautological on 97% of rows, so the model can hit this with role-inference alone.
Intended use
Ranker head of a multi-stage episodic-ingestion compiler. Input is a JSON-serialized (requested field, candidate event) record. Output is a scalar logit that, after grouped softmax over siblings in the same trace, estimates the probability that the candidate event supports the field.
See docs/ranker-hypothesis-log-2026-05-08.md in the repo for the full experimental ladder (H1-H3b-span falsified, label-pivot kept, H3b-replay-v2 falsified, v2-H4-320 kept).
