CoolFace
Modelpublic

abhishekai/gemma-2-2b-legal-sft-v2

sourceHugging Facegemmaupdated 1mo agoView on Hugging Face
0likes28downloads
Model Card

gemma-2-2b-legal-sft-v2

Gemma-2-2B fine-tuned with QLoRA (4-bit) for grounded question answering over legal and financial passages.

It beats the model it was fine-tuned from

score /10
this model8.98 [8.79, 9.17]
base google/gemma-2-2b-it8.40
an earlier legal fine-tune of the same base8.06

Paired difference over base +0.58, bootstrap 95% CI [+0.33, +0.83] — excludes zero. Sign test p = 0.019 over 200 decisive pairs. n=294 of 300 (6 judge API failures).

Fine-tuning a model on a narrow domain usually makes it worse at everything else, and the earlier attempt shows exactly that: 8.06, below the 8.40 base — catastrophic forgetting. The fix was mixing 50% general-domain instruction data back in against the legal half (allenai/tulu-3-sft-mixture, odc-by; 35,540 pairs total). The gain concentrates in instruction following, which is precisely the axis the earlier fine-tune had lost.

How it was trained

QLoRA 4-bit (nf4 + double quantization), LoRA r=16 / α=32 on all attention and MLP projections — 20.8M trainable parameters, 1.28% of the model. 1 epoch, 2,114 steps, effective batch 16, LR 2e-4 cosine to 2e-5. 1×H100, 81 minutes, $5.33.

The adapter is merged into bf16 weights before publishing, so this loads as an ordinary model with no PEFT dependency.

Two details that matter and are easy to get wrong:

  • —Context-free training rows get a general-assistant system prompt, not the "answer only from the provided context" one. Half the training set has no passage; applying the grounded instruction to those rows trains the model on an instruction each row contradicts, which teaches refusal by accident.
  • —The validation split is carved from the training data, never from the judge's eval set. Selecting a checkpoint on the set it will be graded against inflates every later score.

It had not converged. Validation loss was still improving at the final step and early stopping never fired, so 8.98 is a floor for this recipe rather than its ceiling.

Prompt format

Gemma-2's chat template raises on a system role, so the system prompt is folded into the user turn. Its assistant turn ends on `<end_of_turn>` (107), not <eos> — generation must stop on that, and eos_token_id is deliberately the list [1, 107]. Its published config also sets use_cache: false; this repo sets it back to true, without which every decode step re-runs the whole prefix.

How it was evaluated

Claude Sonnet as an independent judge — a different model family from the Gemini 2.5 Flash that wrote the training data, which breaks the circularity that makes a teacher-judged score partly self-congratulation. Four axes out of 10 (question answering, instruction following, grounding in the supplied passage, appropriate refusal), scored against the prompt, the answer, a golden answer and the corpus evidence.

The judge was gated before use: it had to reproduce a known ordering (base Gemma-2-2B above a broken fine-tune of it) before any of its scores were trusted. It did, at p = 0.0022. The base Gemma reference was generated from google/gemma-2-2b-it itself, not a mirror.

Limitations

  • —It answers from a passage you supply. The eval that produced 8.98 hands the model the passage, which favours a retrieval-style model. It says nothing about closed-book performance, where base Gemma-2-2B is stronger.
  • —It invents citations convincingly enough to be dangerous, and has no arithmetic reliability.
  • —Not legal or financial advice. A demonstration of method.
  • —One judge, one eval set, n=294. No labelled benchmark (CaseHOLD, LexGLUE) has been run.
  • —Gemma Terms of Use, not Apache-2.0 — inherited from the base model.