CoolFace
Modelpublic

Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-v2-1-h4-320

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes10downloads
Model Card

ModernBERT field-event ranker (mixed-mode v2.1, H4-320)

Fine-tune of answerdotai/ModernBERT-base with v2.1 labels — same as v2 but with a sharpened `action_causality` heuristic:

  • —Old: tool_result is grounded if ≥2 content tokens appear in ANY later assistant turn. Noise-dominated (file, path, error, exitcode match trivially).
  • —New: require ≥2 distinctive tokens (filter out 50+ corpus-frequent noise tokens) and the assistant turn must be the NEXT one, within 2 events. Selects a proper subset of tool_results based on content grounding.

Key result: per-field improvements on sharpened fields

Fieldv2-320v2.1-320Δ
action_causality0.5610.698+0.137
failed_attempts0.7050.845+0.140
recent_error0.8260.905+0.079
customer_identity0.8110.864+0.053
next_actions0.7550.786+0.031
initiating_command0.9660.968+0.002
outcomeoflatest_attempt0.9660.938-0.028
attempted_actions0.4800.473-0.007
observed_outcomes0.4910.486-0.005

Overall MRR stayed flat at 0.583 (aggregate), but top-1 moved from 0.375 → 0.381. action_causality's eval mass shrank (141 → 50 pairs) because the sharper heuristic only emits the field when the signal is real, which is why the big per-field MRR gain doesn't translate 1:1 into aggregate gain.

Training details

  • —Train rows: 1217
  • —Eval rows: 220
  • —Steps: 320 × accum=8 = 2560 forwards
  • —Last loss: 0.0007
  • —Peak VRAM: 4.08 GiB

Eval metrics

metricvalue
overall MRR0.583
top-10.381
top-20.576
top-30.722
top-50.904
mean expected rank2.69

Per-field MRR (all fields):

fieldnMRRtop-1
attempted_actions8460.4730.24
observed_outcomes8450.4860.24
initiating_command1990.9680.94
outcome_of_latest_attempt1810.9380.88
attempt_outcome_pairs530.6080.34
action_causality500.6980.48
failed_attempts360.8450.72
next_actions320.7860.59
recent_error290.9050.83
customer_identity110.8640.73
transaction_reference50.5830.40
product_name41.0001.00
discarded_options20.2920.00
invalidation_hints20.6250.50
non_promotable_context20.3750.00
payment_or_warranty_detail20.1700.00
assistant_claims_to_verify11.0001.00
explicit_decisions11.0001.00
resolved_context10.3330.00
touched_files10.3330.00
unsupported_hypotheses10.2500.00

Lineage

checkpointMRRtop-1note
mixed-v1 perf-H40.5060.260role-tautological labels
mixed-v2 perf-H4 (160)0.5260.296semantic labels, non-converged
mixed-v2 H4-3200.5830.375converged
mixed-v2.1 H4-320 (this)0.5830.381sharpened causality