cs-552-2026-Clanker-Scientists/adaptor-marianmt-fr-en-legal-v3
04
Adaptor v3 — MarianMT opus-mt-fr-en FT on JRC-Acquis
Component of CS-552 Spring 2026 multi-agent legal QA pipeline (Clanker-Scientists team). Translates French legal evidence (from the FR jurisdiction agent) into English for the Coordinator.
Training
- Base:
Helsinki-NLP/opus-mt-fr-en - Data: JRC-Acquis (European Acquis Communautaire) fr-en parallel sentences, sourced from OPUS object storage
- Filter: 100-400 characters (drops short boilerplate where base saturates; drops long sentences that exceed the 192-token context)
- Train size: 10,000 sentence pairs
- Epochs: 3 at LR 2e-5, beam-6 decoding at inference
Evaluation
Across three filter-stringency iterations:
- v1 (≥20 char): Δ +0.18
- v2 (≥60 char): Δ +0.67
- v3 (100-400 char): Δ +0.98 (monotonic improvement with filter stringency)
Usage
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tok = AutoTokenizer.from_pretrained("cs-552-2026-Clanker-Scientists/adaptor-marianmt-fr-en-legal-v3")
model = AutoModelForSeq2SeqLM.from_pretrained("cs-552-2026-Clanker-Scientists/adaptor-marianmt-fr-en-legal-v3")
fr = "Le contrat est résilié de plein droit en cas de force majeure."
ids = tok(fr, return_tensors="pt").input_ids
out = model.generate(ids, num_beams=6, max_length=192)
print(tok.decode(out[0], skip_special_tokens=True))