CoolFace
Modelpublic

Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-v2-h4-320

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes8downloads
Model Card

episodic-ingestion-compiler / ModernBERT field-event ranker (mixed-mode v2, H4-320)

Fine-tune of answerdotai/ModernBERT-base for the grouped multi-positive softmax ranker task, trained on v2 semantic-reasoning labels for 320 optimizer steps (vs 160 in the initial v2 checkpoint). The extra training let the model fully converge on both legacy and semantic fields.

Versus v2 at 160 steps

Metricv2-H4 (160)v2-H4 (320)Δ
overall MRR0.5260.583+0.057
top-10.2960.375+0.079
last loss0.08540.0000 (converged)

Training details

  • —Train rows: 1217 (mixed-mode v2)
  • —Train groups: 5882
  • —Eval rows: 220 (21986 candidate pairs)
  • —Steps: 320 optimizer steps × accum=8 = 2560 forwards
  • —Optimizer: AdamW lr 5.7e-5, warmup 30
  • —Precision: bf16 autocast + gradient checkpointing
  • —Peak VRAM: 4.08 GiB
  • —Hardware: RTX 5090 Laptop (Blackwell SM 12.0, 24 GiB)

Eval metrics

metricvalue
overall MRR0.583
top-1 recall0.375
top-2 recall0.580
top-3 recall0.737
top-5 recall0.909
mean expected rank2.67

Per-field MRR (top fields in eval):

fieldnMRRtop-1randomratio
attempted_actions8460.4800.240.2531.90x
observed_outcomes8450.4910.240.2531.94x
initiating_command1990.9660.940.2753.51x
outcome_of_latest_attempt1810.9660.940.2763.50x
action_causality1410.5610.300.2542.21x
attempt_outcome_pairs530.6160.340.2652.33x
failed_attempts360.7050.530.2562.75x
next_actions320.7550.560.2962.55x
recent_error290.8260.720.2563.23x
customer_identity110.8110.73--
transaction_reference50.6330.40--
product_name41.0001.00--
discarded_options20.3750.00--
invalidation_hints20.7500.50--
non_promotable_context20.4170.00--
payment_or_warranty_detail20.3330.00--
assistant_claims_to_verify10.5000.00--
explicit_decisions11.0001.00--
resolved_context10.3330.00--
touched_files10.2000.00--
unsupported_hypotheses11.0001.00--

Lineage

checkpointlabelsstepsMRRtop-1notes
V2 (commit 87cb089)conversational1600.6780.4408-row adversarial eval only
mixed-mode stage-4v11600.3230.130pre-perf stack
mixed-mode perf-H4v11600.5060.260perf-H4 stack but v1 (role-tautological) labels
mixed-mode v2-H4v21600.5260.296first v2 train — non-converged
mixed-mode v2-H4-320 (this)v23200.5830.375converged — strongest checkpoint

Semantic-reasoning fields show clear learning

The fields designed to NOT be role-tautological reached high MRR through training, confirming the model is actually reasoning beyond role:

  • —initiating_command: MRR 0.966 (3.51x random) — learns trace boundaries
  • —outcome_of_latest_attempt: MRR 0.966 (3.50x random) — matches call_id chains
  • —recent_error: MRR 0.826 (3.23x random) — latest-failure reasoning
  • —action_causality: MRR 0.561 (2.21x random) — content grounding across turns

Legacy fields (attempted_actions, observed_outcomes) sit near 0.48, close to random for multi-positive softmax. They're role-tautological on 97% of rows, so the model can hit this with role-inference alone.

Intended use

Ranker head of a multi-stage episodic-ingestion compiler. Input is a JSON-serialized (requested field, candidate event) record. Output is a scalar logit that, after grouped softmax over siblings in the same trace, estimates the probability that the candidate event supports the field.

See docs/ranker-hypothesis-log-2026-05-08.md in the repo for the full experimental ladder (H1-H3b-span falsified, label-pivot kept, H3b-replay-v2 falsified, v2-H4-320 kept).