Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-mode-perf-h4
013
episodic-ingestion-compiler / ModernBERT field-event ranker (perf-H4, mixed-mode)
Fine-tune of answerdotai/ModernBERT-base (149 M params, 1 regression output) for the grouped multi-positive softmax ranker task in the episodic-ingestion-compiler training pipeline: given a requested extraction field and a set of candidate source events in a trace, rank each event by how strongly it supports the field.
Training details
- Train set: 1201 rows across 4 tagged modes:
- 109 conversational (original hand-authored)
- 313 autonomous agent loops (generated via local GLM harness)
- 892 prompt→action (mined from local codex/claude session jsonl)
- 99 customer-support (from
jonathansuru/customer_service_information_extraction) - Eval set: 220 rows of held-out mixed-mode data (session-hash disjoint from train)
- Steps: 160 optimizer steps (
--accum-steps 8→ 1280 forwards) - Optimizer: AdamW, lr 5.7e-5 (= 2e-5 × sqrt(8) for effective batch size 8), warmup 30
- Precision: bf16 autocast (no GradScaler)
- Memory optimizations: gradient checkpointing (
use_reentrant=False) - Final train loss: 0.0000
- Peak VRAM: 3.858 GiB (was 12.57 without perf stack)
- Hardware: single RTX 5090 Laptop (Blackwell SM 12.0, 24 GiB), CUDA 13.2, PyTorch 2.11
Eval metrics (220-row mixed eval, 1792 source pairs)
Per-field MRR (top-n fields in the eval):
Lineage vs earlier checkpoints
Intended use
This is the ranker head of a multi-stage episodic-ingestion compiler — NOT a general-purpose classifier. Input is a JSON-serialized (requested field, candidate event) record from the compiler's training pipeline. Output is a scalar logit that, after grouped softmax over siblings in the same trace, estimates the probability that the candidate event supports the requested field.
Limitations
- Current eval shows
attempted_actionsandobserved_outcomes(the two largest field slices, 846 + 845 of 1792 eval pairs) stuck at MRR ≈ 0.49. This is the toolcall-vs-toolresult confusion inherent to per-candidate context-free scoring. Addressing it is planned future work (marker-pool architecture at mixed-mode scale). - Training loss of 0.0001 indicates the model has converged on train. Longer training at this config is unlikely to improve eval.
- Not intended for deployment as a standalone extractor; it's one head in a larger ingestion pipeline.
