CoolFace
Modelpublic

Agreemind/deberta-unfair-tos

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes14downloads
Model Card
⚠️ DEPRECATED — This model used a non-standard training methodology (Focal Loss + class weighting) and was evaluated on a synthetic test set. For accurate, reproducible results on the official LexGLUE benchmark, use our updated models: | Model | μ-F1 | m-F1 | |-------|------|------| | [lexglue-roberta-unfair-tos](https://huggingface.co/Agreemind/lexglue-roberta-unfair-tos) | 96.1 | 84.4 | | lexglue-legalbert-unfair-tos | 96.0 | 84.1 | | lexglue-deberta-unfair-tos | 95.6 | 82.2 | | lexglue-legalbert-small-unfair-tos | 95.0 | 78.5 |

deberta-unfair-tos

Best performing model for UNFAIR-ToS classification

Model Description

This model is fine-tuned on the LexGLUE UNFAIR-ToS dataset to detect unfair clauses in Terms of Service documents.

Base Model: microsoft/deberta-base

Performance

MetricScore
Exact Match Accuracy78.8%
Micro-F10.87
Precision0.98

Risk Categories

The model classifies text into 8 risk categories:

IDCategory
0Limitation of liability
1Unilateral termination
2Unilateral change
3Content removal
4Contract by using
5Choice of law
6Jurisdiction
7Arbitration

Usage

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_name = "Agreemind/deberta-unfair-tos"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

text = "We reserve the right to terminate your account at any time."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)

with torch.no_grad():
    outputs = model(**inputs)
    probs = torch.sigmoid(outputs.logits)

# Get predictions
labels = ["Limitation of liability", "Unilateral termination", "Unilateral change", 
          "Content removal", "Contract by using", "Choice of law", "Jurisdiction", "Arbitration"]
          
for label, prob in zip(labels, probs[0]):
    if prob > 0.5:
        print(f"{label}: {prob:.2%}")

Training

  • —Dataset: LexGLUE UNFAIR-ToS (~5,500 samples)
  • —Loss: Focal Loss with class weighting
  • —Optimizer: AdamW with cosine LR schedule
  • —Epochs: 15 (with early stopping)

Limitations

  • —Arbitration class has lower recall (~38%) due to limited training samples
  • —Optimized for English legal text

Citation

bibtex
@misc{agreemind-unfair-tos,
  author = {Agreemind},
  title = {deberta-unfair-tos},
  year = {2024},
  publisher = {HuggingFace},
  url = {https://huggingface.co/Agreemind/deberta-unfair-tos}
}