CoolFace
Modelpublic

Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-mode-perf-h4

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes13downloads
Model Card

episodic-ingestion-compiler / ModernBERT field-event ranker (perf-H4, mixed-mode)

Fine-tune of answerdotai/ModernBERT-base (149 M params, 1 regression output) for the grouped multi-positive softmax ranker task in the episodic-ingestion-compiler training pipeline: given a requested extraction field and a set of candidate source events in a trace, rank each event by how strongly it supports the field.

Training details

  • —Train set: 1201 rows across 4 tagged modes:
  • —109 conversational (original hand-authored)
  • —313 autonomous agent loops (generated via local GLM harness)
  • —892 prompt→action (mined from local codex/claude session jsonl)
  • —99 customer-support (from jonathansuru/customer_service_information_extraction)
  • —Eval set: 220 rows of held-out mixed-mode data (session-hash disjoint from train)
  • —Steps: 160 optimizer steps (--accum-steps 8 → 1280 forwards)
  • —Optimizer: AdamW, lr 5.7e-5 (= 2e-5 × sqrt(8) for effective batch size 8), warmup 30
  • —Precision: bf16 autocast (no GradScaler)
  • —Memory optimizations: gradient checkpointing (use_reentrant=False)
  • —Final train loss: 0.0000
  • —Peak VRAM: 3.858 GiB (was 12.57 without perf stack)
  • —Hardware: single RTX 5090 Laptop (Blackwell SM 12.0, 24 GiB), CUDA 13.2, PyTorch 2.11

Eval metrics (220-row mixed eval, 1792 source pairs)

metricvalue
top-1 recall0.257
top-2 recall0.491
top-3 recall0.677
top-5 recall0.901
MRR0.504
mean expected rank2.92

Per-field MRR (top-n fields in the eval):

fieldnMRRtop-1
attempted_actions8460.4910.24
observed_outcomes8450.4910.24
failed_attempts360.8660.75
next_actions320.7010.47
customer_identity110.6890.55
transaction_reference50.5670.20
product_name40.6250.25
discarded_options20.2920.00
invalidation_hints20.6250.50
non_promotable_context20.6250.50
payment_or_warranty_detail20.3120.00
assistant_claims_to_verify10.2000.00
explicit_decisions10.5000.00
resolved_context10.3330.00
touched_files11.0001.00
unsupported_hypotheses10.2500.00

Lineage vs earlier checkpoints

checkpointtrain rowsMRRtop-1notes
V2 (commit 87cb089)109 conversational0.6780.440Evaluated on 8-row conversational adversarial eval — different and smaller test set
Stage-4 mixed-mode1201 mixed0.3230.130Same 220-row eval as this checkpoint
perf-H4 (this)1201 mixed0.5060.260bf16 + gradient checkpointing + grad accumulation + sqrt LR

Intended use

This is the ranker head of a multi-stage episodic-ingestion compiler — NOT a general-purpose classifier. Input is a JSON-serialized (requested field, candidate event) record from the compiler's training pipeline. Output is a scalar logit that, after grouped softmax over siblings in the same trace, estimates the probability that the candidate event supports the requested field.

Limitations

  • —Current eval shows attempted_actions and observed_outcomes (the two largest field slices, 846 + 845 of 1792 eval pairs) stuck at MRR ≈ 0.49. This is the toolcall-vs-toolresult confusion inherent to per-candidate context-free scoring. Addressing it is planned future work (marker-pool architecture at mixed-mode scale).
  • —Training loss of 0.0001 indicates the model has converged on train. Longer training at this config is unlikely to improve eval.
  • —Not intended for deployment as a standalone extractor; it's one head in a larger ingestion pipeline.