AmirMohseni/modernbert-large-v3-legal-guidance-prefix-len4096-seed42
ModernBERT-large for prefix-aware legal-guidance detection
This model detects whether legal guidance is already present in the cumulative user-turn prefix of a multi-turn LLM conversation. It was trained for research on early detection and routing of legal information needs. It is not a legal advice system and must not be used to determine whether a person has a valid legal claim.
Task formulation
One example is generated after each user turn. Prefixes before first_guidance_user_turn_id are negative; the onset prefix and all later prefixes are positive. Every prefix from a guidance-negative conversation is negative. The final prefix matches the user-only full-conversation serialization used in the comparison model.
All prefixes are retained. Each prefix is weighted by the inverse of its conversation's number of user turns, so every conversation has total training weight 1.0.
Data
- Dataset: AmirMohseni/WildChat-Legal-Classification-V3-Hierarchical
- Requested revision:
main(latest at run time) - Resolved revision:
489b87c18e571e8c7e9a3b1c2463b79e50cef120 - Training: 1,632 conversations / 4,865 prefixes
- Validation: 290 conversations / 912 prefixes
- Positive prefixes: 2,042 train / 376 validation
The first-guidance-turn annotations are silver labels and were not independently human-validated for this run. Onset metrics must therefore not be described as gold evaluation.
Training configuration
Validation results
The threshold 0.36 was selected on the silver validation split by conversation-weighted prefix macro-F1, with positive-class F1 as tie-breaker.
Silver onset analysis
Repository artifacts
prefix_validation_metrics.json: complete metrics and threshold curveprefix_validation_predictions.csv: per-prefix probabilities and predictionsonset_validation_predictions.csv: conversation-level onset predictionsrun_metadata.json: model, data, software, and hardware provenance
Limitations
The dataset is English-language, jurisdiction-agnostic, and derived from public LLM interaction logs that may contain sensitive content. The training and onset labels are silver labels. The model was trained with one seed and has not yet been evaluated on the adjudicated 200-conversation human gold set. Final-prefix guidance evaluation on gold is appropriate after adjudication; onset evaluation on gold requires separate human review of first_guidance_user_turn_id.
