CoolFace
Modelpublic

pyloxsystems/legal-cuad-llama-3.1-8b-lora

sourceHugging Facellama3.1updated 5mo agoView on Hugging Face
0likes14downloads
Model Card

Pylox Legal Contract Q&A 8B (legal-cuad)

A LoRA adapter for meta-llama/Llama-3.1-8B-Instruct, fine-tuned on the CUAD (Contract Understanding Atticus Dataset) clause Q&A pairs. Built end-to-end on a single NVIDIA Grace Blackwell GB10 (DGX Spark, 128 GB unified memory) with the same NF4 train, NVFP4 serve, EAGLE-3 speculative decoding stack used across the Pylox Forge portfolio.

This is a small-dataset seed adapter intended as a portfolio demonstration of the legal vertical. The fine-tune-quality measurement against the base model is below the 50/50 line on this volume of data (see Evaluation), and is documented honestly. A larger refresh round on a real contract corpus is the recommended path to production-grade clause extraction.

Model details

  • —Adapter: LoRA (PEFT, rank 32, alpha 64, dropout 0.1)
  • —Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • —Base model: meta-llama/Llama-3.1-8B-Instruct
  • —Recommended serving base: nvidia/Llama-3.1-8B-Instruct-NVFP4
  • —Speculative head: RedHatAI/Llama-3.1-8B-Instruct-speculator.eagle3
  • —License: Llama 3.1 Community License
  • —Hardware: NVIDIA Grace Blackwell GB10 (DGX Spark, 128 GB UMA)

Training data and technique

  • —Source: theatticusproject/cuad-qa (Contract Understanding Atticus Dataset)
  • —Source dataset size: 510 commercial contract clause Q&A pairs (CUAD reference)
  • —Examples after dedup, quality filter, and PII redaction: 321
  • —Format: Standard chat messages format with assistant-only loss
  • —Method: NF4 QLoRA SFT (4-bit NormalFloat base, bfloat16 compute, double quantization)
  • —Hyperparameters: 3 epochs, cosine LR (2e-4 peak), maxseqlength 2048, sequence packing enabled, NEFTune noise alpha 5, pagedadamw8bit optimizer, gradient accumulation 16

Evaluation

Numbers are pulled directly from the local benchmark JSON. No invented values.

Throughput on Grace Blackwell GB10

MetricValue
Single-user throughput40.7 tok/s
Concurrent batch-8 throughput166.9 tok/s
TTFT p50130 ms
End-to-end latency p503705 ms

Capability preservation (academic, lm-evaluation-harness, 500 samples each)

BenchmarkScore
HellaSwag (common sense)76.0%
TruthfulQA (hallucination resistance)30.0%
MMLU-Pronot run

Fine-tune quality vs base (LLM judge, pairwise)

MetricValue
Win rate vs meta-llama/Llama-3.1-8B-Instruct15.0%

Honest read: at 321 training examples the adapter does not beat the base model in pairwise judge quality. CUAD-QA's clause taxonomy is also narrow relative to real-world contract diversity. The pipeline runs end-to-end and the artifacts ship, but production-grade clause extraction across a real contract repository requires a substantially larger and more diverse corpus. A LegalBench / CUAD F1 evaluation has not been run; that scoring is the next step before this adapter is used for any real review.

Safety (with Pylox safety gateway, 50-prompt red team)

MetricValue
Adversarial block rate75.56%
False positive rate on benign controls0.0%

Zero false positives on benign controls is the right safety profile for a legal-domain adapter (over-blocking on contract questions is itself a quality failure). Adversarial block rate at 75.56% is below frontier-aligned models but acceptable for a portfolio piece routed through a safety gateway.

Quickstart

PEFT (research / batch)

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

base_id = "meta-llama/Llama-3.1-8B-Instruct"
adapter_id = "pyloxsystems/legal-cuad-llama-3.1-8b-lora"

tokenizer = AutoTokenizer.from_pretrained(base_id)
model = AutoModelForCausalLM.from_pretrained(
    base_id, torch_dtype=torch.bfloat16, device_map="auto"
)
model = PeftModel.from_pretrained(model, adapter_id)

messages = [
    {"role": "system", "content": "You are a contract review assistant. Cite the clause and explain plainly."},
    {"role": "user", "content": "Does this NDA include a non-solicitation clause?\n\n[paste contract text]"},
]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))

vLLM with NVFP4 base and EAGLE-3 speculative decoding

bash
vllm serve nvidia/Llama-3.1-8B-Instruct-NVFP4 \
    --enable-lora \
    --lora-modules legal-cuad=pyloxsystems/legal-cuad-llama-3.1-8b-lora \
    --speculative-config '{
        "method": "eagle3",
        "model": "RedHatAI/Llama-3.1-8B-Instruct-speculator.eagle3",
        "num_speculative_tokens": 5
    }'

Intended use

  • —Clause-by-clause Q&A on US-English commercial contracts
  • —Due-diligence reading of NDAs, MSAs, and similar agreements under 20 pages
  • —Risk-flag extraction from contracts in CUAD's taxonomy (assignment, change of control, indemnity, etc.)
  • —Pipeline demonstration for evaluating the Pylox Forge stack on a legal vertical

Out of scope

  • —Final legal advice. All outputs require licensed-attorney review before any decision.
  • —Contracts longer than 20 pages without an external retrieval layer
  • —Court filings, statutes, or case law (this adapter was trained on contracts only)
  • —Non-English contracts (French, Spanish, German performance unmeasured)
  • —Production legal review without a substantially larger fine-tune corpus

Limitations

  • —321 training examples is small relative to production-grade legal fine-tunes. Fine-tune quality vs base is below the 50/50 line.
  • —CUAD's taxonomy is narrow; the adapter may echo CUAD clause categories even when client uses different terminology.
  • —LegalBench / CUAD F1 not yet measured. Recommended next step.
  • —Sequence length capped at 2048; long contracts must be chunked externally.

License

Inherits the Llama 3.1 Community License from the base model.

Citation

bibtex
@misc{pylox_legal_cuad_2026,
  author       = {Girard, Emilio},
  title        = {Pylox Legal Contract Q\&A 8B (legal-cuad)},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/pyloxsystems/legal-cuad-llama-3.1-8b-lora}}
}

Pylox Forge is a solo-operated LLM fine-tuning lab on NVIDIA Grace Blackwell. Site: pyloxforge.com. Other adapters: pyloxsystems on Hugging Face.