CoolFace
Modelpublic

cs-552-2026-Clanker-Scientists/adaptor-marianmt-fr-en-legal-v3

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes4downloads
Model Card

Adaptor v3 — MarianMT opus-mt-fr-en FT on JRC-Acquis

Component of CS-552 Spring 2026 multi-agent legal QA pipeline (Clanker-Scientists team). Translates French legal evidence (from the FR jurisdiction agent) into English for the Coordinator.

Training

  • —Base: Helsinki-NLP/opus-mt-fr-en
  • —Data: JRC-Acquis (European Acquis Communautaire) fr-en parallel sentences, sourced from OPUS object storage
  • —Filter: 100-400 characters (drops short boilerplate where base saturates; drops long sentences that exceed the 192-token context)
  • —Train size: 10,000 sentence pairs
  • —Epochs: 3 at LR 2e-5, beam-6 decoding at inference

Evaluation

Test setMetricBaseFTΔ
JRC-Acquis fr-en, 100-400 char, N=200corpus BLEU57.6958.67+0.98

Across three filter-stringency iterations:

  • —v1 (≥20 char): Δ +0.18
  • —v2 (≥60 char): Δ +0.67
  • —v3 (100-400 char): Δ +0.98 (monotonic improvement with filter stringency)

Usage

python
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tok = AutoTokenizer.from_pretrained("cs-552-2026-Clanker-Scientists/adaptor-marianmt-fr-en-legal-v3")
model = AutoModelForSeq2SeqLM.from_pretrained("cs-552-2026-Clanker-Scientists/adaptor-marianmt-fr-en-legal-v3")

fr = "Le contrat est résilié de plein droit en cas de force majeure."
ids = tok(fr, return_tensors="pt").input_ids
out = model.generate(ids, num_beams=6, max_length=192)
print(tok.decode(out[0], skip_special_tokens=True))