CoolFace
Modelpublic

saidutta69/PhishScout

sourceHugging Facemitupdated 15d agoView on Hugging Face
0likes
Model Card

PhishScout

<div align="center"> <img src="https://photu.kashyalabanavli.site/racer-is-op.png" alt="RACER IS OP" width="100%"> </div>

<br>

Catch phishing URLs before they catch you — 131.5 KB of on-device LightGBM. No GPU, no cloud, no API calls. 35 URL features, trained on PhishTrap's auto-refreshing 20K-URL dataset.

Priorities: Quality > Size > Speed

Trained on: saidutta69/PhishTrap — auto-refreshed every 6 hours from OpenPhish, Phishing.Database, PhishStats, and Tranco


Model Overview

PhishScout is a tiny gradient-boosted tree model that classifies URLs as phishing or legitimate using URL structure alone — no page content, no third-party lookups, no network calls at inference time. The full model is a single 131.5 KB ONNX file that runs anywhere ONNX Runtime runs (Python, Node, browser via ONNX Runtime Web, mobile).

PropertyValue
ArchitectureLightGBM (gradient boosting)
Config60 trees, max_depth 8, 31 leaves
Features35 (URL-structural + semantic signals)
Model size131.5 KB (ONNX, opset 15)
Training time~0.5s on CPU
Inference latency~9 µs/URL (median, CPU)
ONNX parity100% (predictions match LightGBM exactly)
LicenseMIT

Use Cases

Primary — Real-time, on-device URL screening

  • —Browser extension — screen every URL before navigation/click, entirely locally. Zero data leaves the device; works offline; no privacy trade-offs.
  • —Email security clients — pre-click link inspection in desktop/mobile mail apps; flag or block suspicious links before the user opens them.
  • —SMS / messaging link scanners — instant check of short links from unknown senders in WhatsApp, Telegram, and SMS apps.
  • —QR code scanner apps — decode -> classify -> warn, before opening the URL.

Secondary — Infrastructure integration

  • —Email gateway & firewall plugins — lightweight first-pass filter; hand off high-probability URLs to heavier sandboxing for deep inspection.
  • —Proxy / DNS-layer filtering — classify requested URLs at the edge with single-digit-microsecond latency.
  • —SIEM / SOAR enrichment — batch-classify URLs from logs, tickets, and incident data to prioritize triage.
  • —Security awareness training — generate realistic phishing examples or explain why a URL is suspicious (feature diagnostics).

Tertiary — Research & data pipelines

  • —URL crawler triage — pre-filter the streams of phishing feeds, mirroring how the PhishTrap dataset is built.
  • —Phishing trend measurement — lightweight prevalence estimation across URL corpora.
  • —Benchmark baseline — a compact, reproducible reference model for comparing larger detectors.

Performance

Final model (test split, held out, seed 42)

MetricValue
Accuracy0.8697
Precision0.9033
Recall0.8254
F10.8626
ROC AUC0.9369

Confusion matrix (test, 2,993 URLs): TP=1224, FP=131, FN=259, TN=1379.

Determinism & stability

  • —10-seed variance on test: F1 0.8626 +/- 0.0000, AUC 0.9369 +/- 0.0000 — fully deterministic.

Robustness

PerturbationF1
Clean0.8626
5% extreme-value corruption0.8080
Gaussian noise sigma=0.50.7097

Threshold behavior

  • —Default threshold 0.5 is near-optimal.
  • —F1-optimal threshold tuned on val: t=0.36 -> test F1 0.8529 (trades precision for recall: rec 0.869, prec 0.818).

Model Selection — Full Benchmark

PhishScout is the winner of a 34-config Pareto sweep plus a 12-model baseline benchmark. Every candidate: 5-fold CV on train, held-out test eval, ONNX export, size, and parity.

Baseline candidates (16 features)

modelCV F1test F1accprecrecAUCsize KB
gaussian_nb0.55710.5420.6810.9390.3810.8631.8
logistic_regression0.8160.8160.8250.8540.7810.9010.7
decision_tree d80.84010.8460.8530.8840.8110.91914.0
random_forest 10x60.82460.8290.8400.8790.7840.90732.4
extra_trees 10x60.79530.8010.8120.8440.7610.88014.9
adaboost 30x30.83050.8260.8340.8610.7940.90333.7
gradient_boosting 30x30.82830.8320.8410.8710.7960.90916.8
lightgbm 30x6L150.84150.8460.8540.8900.8050.92131.9
xgboost 30x40.83670.8400.8480.8740.8080.91428.0
svm rbf0.83240.8360.8470.8920.7870.914444.4
knn k=50.83060.8330.8410.8650.8050.897889.4

Pareto sweep (16 vs 35 features, 34 configs)

Pareto frontier (F1 vs ONNX size):

modelfeatsF1accAUCsize KB
dt_d8full0.84460.85270.920811.5
dt_d8base160.84560.85330.919014.0
xgb_30x4full0.84720.85470.923928.1
lgbm_30x6L15full0.85020.85670.927031.9
gb_50x4full0.85060.85730.929453.9
lgbm_60x8L31base160.85450.86200.9285131.5
PhishScout (lgbm_60x8L31)full0.86260.86970.9369131.5

Key findings:

  • —Everything above 131.5 KB does worse (overfitting): lgbm100x10L63 0.8589 @ 447 KB, rf100x12 0.8566 @ 2388 KB, lgbm_200x12L127 0.8477 @ 1894 KB. 131.5 KB is the unambiguous elbow.
  • —The 35-feature set beats 16 features at every size: +0.018 F1 for the 131.5 KB LightGBM config.
  • —The expanded features (keywords, brand impersonation, punycode, ports, hex-IP, encoding tricks) are what unlock the quality jump — all 16 base features verified byte-exact against the stored dataset.

Why not the smaller models?

The 14 KB decision tree (F1 0.846) is the best sub-50 KB option, but PhishScout beats it statistically significantly: McNemar p=0.0002 (lgbm-only-right=115 vs tree-only-right=64), paired bootstrap 95% CI of F1 difference [+0.0084, +0.0271] — entirely positive. +0.018 F1 / +0.017 AUC for 9x the size of a file that is 0.1% of any app.


Features (35)

Base 16 — URL structure

FeatureDescription
url_lengthTotal URL character count
hyphen_countNumber of hyphens
digit_countNumber of digits
subdomain_countNumber of subdomains
trusted_tld1 if TLD is .com/.org/.net/.edu/.gov
protocol_exists1 if http/https present
special_char_countSpecial characters (@, -, _, ., etc.)
entropyShannon entropy of URL string
path_depthNumber of path segments
domain_lengthLength of domain portion
is_domain_ip1 if domain is an IP address
has_at_symbol1 if @ present
has_double_slash_redirect1 if // appears after protocol
tld_lengthLength of top-level domain
query_param_countNumber of query parameters
path_lengthLength of path portion

Expanded 19 — Semantic signals

FeatureDescription
keyword_countOccurrences of 28 suspicious keywords (login, verify, secure, wallet, refund, ...)
brand_in_host1 if one of 33 well-known brands in hostname
brand_in_path1 if a brand in the path
brand_countTotal brand mentions
has_punycode1 if hostname uses xn-- punycode
has_nonstandard_port1 if explicit port other than 80/443
percent_encoding_countCount of %-escapes
has_hex_ip1 if hex-encoded IP (0x...) present
uppercase_ratioFraction of uppercase characters
www_countOccurrences of "www."
max_digit_runLongest consecutive digit run
max_label_lenLongest hostname label
host_label_countNumber of hostname labels
has_domain_in_path1 if a domain-like token appears in the path
query_lengthQuery string length
equals_countNumber of = in query
sensitive_extension1 if ends with .exe/.scr/.apk/.zip/.jar/...
has_multiple_tlds1 if multiple TLD-like tokens with slashes
consecutive_punct1 if 3+ punctuation chars in a row

Usage

Python (ONNX Runtime)

python
import onnxruntime as ort
import numpy as np

sess = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])
features = np.array([[ /* 35 features, order below */ ]], dtype=np.float32)
pred, proba = sess.run(None, {"input": features})
# pred == 1 -> phishing, pred == 0 -> legitimate

Feature order (35):

url_length, hyphen_count, digit_count, subdomain_count, protocol_exists,
special_char_count, entropy, path_depth, domain_length, is_domain_ip,
has_at_symbol, has_double_slash_redirect, tld_length, trusted_tld,
query_param_count, path_length, keyword_count, brand_in_host, brand_in_path,
brand_count, has_punycode, has_nonstandard_port, percent_encoding_count,
has_hex_ip, uppercase_ratio, www_count, max_digit_run, max_label_len,
host_label_count, has_domain_in_path, query_length, equals_count,
sensitive_extension, has_multiple_tlds, consecutive_punct

Browser (ONNX Runtime Web)

Load model.onnx into ort.InferenceSession with WebAssembly backend — runs fully in-browser, no server.


Verification & Reproducibility

  • —ONNX parity — ONNX Runtime predictions match LightGBM on 100% of 2,000 sampled test URLs.
  • —Deterministic — identical results across 10 seeds (LightGBM with fixed seed).
  • —No leakage — train/val/test splits are disjoint (verified in the dataset build); test set untouched during model selection except for the final evaluation.
  • —Reproducible — fixed seed 42; full training and benchmark code open source at github.com/instax-dutta/PhishScout (train_phishscout.py, benchmark_models.py, sweep_models.py, deep_benchmark.py).

Limitations

  • —URL-only signals — no page content, TLS certificate, or behavioral analysis. A URL that looks benign but hosts phishing content can be missed.
  • —Bare trusted-domain phishing — the hardest case for any URL-only model: phishing hosted on the actual root of a trusted domain (e.g. paypal.com/... after compromise) is largely invisible to structure-based features. 63-69% of false negatives are URLs with trusted TLDs and no subdomains.
  • —English-centric keyword/brand lists — the 28 keywords and 33 brands are English/Western-market oriented; detection quality degrades for other languages and smaller regional brands.
  • —Feed dependence — training data quality is inherited from PhishTrap's live feeds; retrain on the latest dataset snapshot for current campaigns (pipeline refreshes every 6 hours).
  • —No temporal adaptation on its own — the model does not adapt between retrains; novel phishing techniques appear between dataset refreshes.
  • —Threshold sensitivity — at default 0.5 the model favors precision (0.903); use t=0.36 for higher recall at lower precision, or calibrate per use case.

Training Data

Trained on PhishTrap v2 (build 20260731T085613Z): 19,954 URLs, balanced 50/50, 70/15/15 split. Phishing URLs from PyFunceble-verified feeds (Phishing.Database, OpenPhish, PhishStats); legitimate URLs from Tranco top 10K. See the dataset card for full provenance.


Citation

bibtex
@misc{saidutta69_2026_phishscout,
  author = {Sai Dutta Abhishek Dash},
  title = {PhishScout: Tiny On-Device Phishing URL Detector},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/models/saidutta69/PhishScout}},
  note = {Trained on the auto-refreshing PhishTrap dataset}
}

Built on PhishTrap + Phishing.Database, OpenPhish, PhishStats, and Tranco. MIT licensed.