CoolFace
Modelpublic

perfecXion/intentguard-finance

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes9downloads
Model Card

IntentGuard — Financial Services

![License](https://opensource.org/licenses/Apache-2.0) ![Accuracy](#performance) ![Size](#model-details) ![Latency](#performance) ![Format](#model-details)

Production-ready vertical intent classifier for LLM chatbot guardrails. Classifies user messages as `allow`, `deny`, or `abstain` to keep financial services chatbots on-topic and secure.

Research Article | perfecXion.ai | Finance Model | Healthcare Model | Legal Model


IntentGuard Model Family

IntentGuard provides specialized intent classifiers for high-stakes verticals where chatbot misuse carries regulatory, legal, or safety risk:

ModelVerticalAccuracyOff-Topic Pass RateLink
intentguard-financeFinancial Services99.6%0.00%This model
intentguard-healthcareHealthcare & Clinical98.9%0.98%perfecXion/intentguard-healthcare
intentguard-legalLegal & Compliance97.9%0.50%perfecXion/intentguard-legal

Overview

The Problem

Enterprise chatbots in regulated industries face a critical challenge: users inevitably ask off-topic questions (sports, entertainment, relationship advice) that the underlying LLM will happily answer — exposing the organization to compliance risk, brand damage, and potential liability.

Traditional keyword filters miss nuanced off-topic queries, while LLM-based guardrails are too slow and expensive for real-time inference.

The Solution

IntentGuard uses a tiny, purpose-trained DeBERTa-v3-xsmall model (22M parameters, 2.5MB quantized) to classify user intent in <30ms on CPU. The three-way classification (allow/deny/abstain) enables precise control:

  • Allow — On-topic for the vertical, pass to the LLM
  • Deny — Clearly off-topic, block with a polite redirect
  • Abstain — Ambiguous, escalate to secondary classifier or human review

Performance

MetricValue
Overall Accuracy99.6%
Legitimate Block Rate0.00% (no false positives)
Off-Topic Pass Rate0.00% (no false negatives)
p99 Latency (CPU)<30ms
Model Size (ONNX INT8)2.5MB
Base Parameters22M (DeBERTa-v3-xsmall)
Expected Calibration Error<0.03

Classification Decision Framework

User Message → Tokenize → DeBERTa Inference → Softmax
                                                  ↓
                                    ┌──────────────┼──────────────┐
                                    │              │              │
                                  ALLOW          DENY         ABSTAIN
                                (on-topic)    (off-topic)    (uncertain)
                                    │              │              │
                                Pass to LLM   Block + Redirect  Escalate

Model Details

PropertyValue
ArchitectureDeBERTa-v3-xsmall (fine-tuned for 3-way classification)
FormatONNX (INT8 quantized)
Version1.0
VerticalFinance (Financial Services)
TrainingSupervised fine-tuning on curated intent datasets
QuantizationINT8 via ONNX Runtime
GPU RequiredNo — runs on CPU
PublisherperfecXion.ai

Core Topics (Allow)

Banking, lending, credit, payments, investing, insurance, tax, personal finance, retirement, mortgages, financial planning, budgeting

Hard Exclusions (Deny)

Sports, entertainment, cooking, gaming, celebrity gossip, fashion, travel/leisure, fiction writing, relationship advice


Usage

Python (ONNX Runtime)

python
import onnxruntime as ort
from transformers import AutoTokenizer
import numpy as np

# Load model and tokenizer
tokenizer = AutoTokenizer.from_pretrained("perfecXion/intentguard-finance")
session = ort.InferenceSession("model.onnx")

# Classify a user message
text = "What are the current mortgage rates for a 30-year fixed loan?"
inputs = tokenizer(text, return_tensors="np", max_length=128, truncation=True, padding="max_length")

logits = session.run(None, {
    "input_ids": inputs["input_ids"],
    "attention_mask": inputs["attention_mask"]
})[0]

labels = ["allow", "deny", "abstain"]
prediction = labels[np.argmax(logits)]
confidence = float(np.max(np.exp(logits) / np.sum(np.exp(logits))))

print(f"Intent: {prediction} (confidence: {confidence:.3f})")
# Output: Intent: allow (confidence: 0.998)

Docker

bash
# Pull and run the container
docker pull ghcr.io/perfecxion/intentguard:finance-1.0
docker run -p 8080:8080 ghcr.io/perfecxion/intentguard:finance-1.0

# Classify a message
curl -X POST http://localhost:8080/v1/classify \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "What are the current mortgage rates?"}]}'

# Response: {"intent": "allow", "confidence": 0.998}

pip

bash
pip install intentguard

# Python usage
from intentguard import IntentGuard

guard = IntentGuard.load("finance")
result = guard.classify("What are the current mortgage rates?")
print(result)  # Intent(label='allow', confidence=0.998)

Example Classifications

User MessagePredictedConfidenceCorrect?
"What are mortgage rates for a 30-year fixed?"allow0.998
"How do I open a Roth IRA?"allow0.997
"Who won the Super Bowl?"deny0.999
"Tell me a joke"deny0.996
"Is my health insurance FSA-eligible?"allow0.942✅ (financial context)
"What's the weather today?"deny0.998

Citation

bibtex
@misc{thornton2025intentguard,
  title={IntentGuard: A Production-Grade Vertical Intent Classifier for LLM Guardrails},
  author={Thornton, Scott},
  year={2025},
  publisher={perfecXion.ai},
  url={https://perfecxion.ai/articles/intentguard-vertical-intent-classifier-llm-guardrails.html},
  note={Model: https://huggingface.co/perfecXion/intentguard-finance}
}

Quality Metrics

MetricResult
Accuracy (Finance vertical)99.6%
Legitimate Block Rate0.00%
Off-Topic Pass Rate0.00%
Expected Calibration Error<0.03
ONNX INT8 QuantizationValidated
CPU Inference (p99)<30ms
Docker ContainerAvailable

License

Apache 2.0


Links