saidutta69/PhishScout
PhishScout
<div align="center"> <img src="https://photu.kashyalabanavli.site/racer-is-op.png" alt="RACER IS OP" width="100%"> </div>
<br>
Catch phishing URLs before they catch you — 131.5 KB of on-device LightGBM. No GPU, no cloud, no API calls. 35 URL features, trained on PhishTrap's auto-refreshing 20K-URL dataset.
Priorities: Quality > Size > Speed
Trained on: saidutta69/PhishTrap — auto-refreshed every 6 hours from OpenPhish, Phishing.Database, PhishStats, and Tranco
Model Overview
PhishScout is a tiny gradient-boosted tree model that classifies URLs as phishing or legitimate using URL structure alone — no page content, no third-party lookups, no network calls at inference time. The full model is a single 131.5 KB ONNX file that runs anywhere ONNX Runtime runs (Python, Node, browser via ONNX Runtime Web, mobile).
Use Cases
Primary — Real-time, on-device URL screening
- Browser extension — screen every URL before navigation/click, entirely locally. Zero data leaves the device; works offline; no privacy trade-offs.
- Email security clients — pre-click link inspection in desktop/mobile mail apps; flag or block suspicious links before the user opens them.
- SMS / messaging link scanners — instant check of short links from unknown senders in WhatsApp, Telegram, and SMS apps.
- QR code scanner apps — decode -> classify -> warn, before opening the URL.
Secondary — Infrastructure integration
- Email gateway & firewall plugins — lightweight first-pass filter; hand off high-probability URLs to heavier sandboxing for deep inspection.
- Proxy / DNS-layer filtering — classify requested URLs at the edge with single-digit-microsecond latency.
- SIEM / SOAR enrichment — batch-classify URLs from logs, tickets, and incident data to prioritize triage.
- Security awareness training — generate realistic phishing examples or explain why a URL is suspicious (feature diagnostics).
Tertiary — Research & data pipelines
- URL crawler triage — pre-filter the streams of phishing feeds, mirroring how the PhishTrap dataset is built.
- Phishing trend measurement — lightweight prevalence estimation across URL corpora.
- Benchmark baseline — a compact, reproducible reference model for comparing larger detectors.
Performance
Final model (test split, held out, seed 42)
Confusion matrix (test, 2,993 URLs): TP=1224, FP=131, FN=259, TN=1379.
Determinism & stability
- 10-seed variance on test: F1 0.8626 +/- 0.0000, AUC 0.9369 +/- 0.0000 — fully deterministic.
Robustness
Threshold behavior
- Default threshold 0.5 is near-optimal.
- F1-optimal threshold tuned on val: t=0.36 -> test F1 0.8529 (trades precision for recall: rec 0.869, prec 0.818).
Model Selection — Full Benchmark
PhishScout is the winner of a 34-config Pareto sweep plus a 12-model baseline benchmark. Every candidate: 5-fold CV on train, held-out test eval, ONNX export, size, and parity.
Baseline candidates (16 features)
Pareto sweep (16 vs 35 features, 34 configs)
Pareto frontier (F1 vs ONNX size):
Key findings:
- Everything above 131.5 KB does worse (overfitting): lgbm100x10L63 0.8589 @ 447 KB, rf100x12 0.8566 @ 2388 KB, lgbm_200x12L127 0.8477 @ 1894 KB. 131.5 KB is the unambiguous elbow.
- The 35-feature set beats 16 features at every size: +0.018 F1 for the 131.5 KB LightGBM config.
- The expanded features (keywords, brand impersonation, punycode, ports, hex-IP, encoding tricks) are what unlock the quality jump — all 16 base features verified byte-exact against the stored dataset.
Why not the smaller models?
The 14 KB decision tree (F1 0.846) is the best sub-50 KB option, but PhishScout beats it statistically significantly: McNemar p=0.0002 (lgbm-only-right=115 vs tree-only-right=64), paired bootstrap 95% CI of F1 difference [+0.0084, +0.0271] — entirely positive. +0.018 F1 / +0.017 AUC for 9x the size of a file that is 0.1% of any app.
Features (35)
Base 16 — URL structure
Expanded 19 — Semantic signals
Usage
Python (ONNX Runtime)
import onnxruntime as ort
import numpy as np
sess = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])
features = np.array([[ /* 35 features, order below */ ]], dtype=np.float32)
pred, proba = sess.run(None, {"input": features})
# pred == 1 -> phishing, pred == 0 -> legitimateFeature order (35):
url_length, hyphen_count, digit_count, subdomain_count, protocol_exists,
special_char_count, entropy, path_depth, domain_length, is_domain_ip,
has_at_symbol, has_double_slash_redirect, tld_length, trusted_tld,
query_param_count, path_length, keyword_count, brand_in_host, brand_in_path,
brand_count, has_punycode, has_nonstandard_port, percent_encoding_count,
has_hex_ip, uppercase_ratio, www_count, max_digit_run, max_label_len,
host_label_count, has_domain_in_path, query_length, equals_count,
sensitive_extension, has_multiple_tlds, consecutive_punctBrowser (ONNX Runtime Web)
Load model.onnx into ort.InferenceSession with WebAssembly backend — runs fully in-browser, no server.
Verification & Reproducibility
- ONNX parity — ONNX Runtime predictions match LightGBM on 100% of 2,000 sampled test URLs.
- Deterministic — identical results across 10 seeds (LightGBM with fixed seed).
- No leakage — train/val/test splits are disjoint (verified in the dataset build); test set untouched during model selection except for the final evaluation.
- Reproducible — fixed seed 42; full training and benchmark code open source at github.com/instax-dutta/PhishScout (
train_phishscout.py,benchmark_models.py,sweep_models.py,deep_benchmark.py).
Limitations
- URL-only signals — no page content, TLS certificate, or behavioral analysis. A URL that looks benign but hosts phishing content can be missed.
- Bare trusted-domain phishing — the hardest case for any URL-only model: phishing hosted on the actual root of a trusted domain (e.g.
paypal.com/...after compromise) is largely invisible to structure-based features. 63-69% of false negatives are URLs with trusted TLDs and no subdomains. - English-centric keyword/brand lists — the 28 keywords and 33 brands are English/Western-market oriented; detection quality degrades for other languages and smaller regional brands.
- Feed dependence — training data quality is inherited from PhishTrap's live feeds; retrain on the latest dataset snapshot for current campaigns (pipeline refreshes every 6 hours).
- No temporal adaptation on its own — the model does not adapt between retrains; novel phishing techniques appear between dataset refreshes.
- Threshold sensitivity — at default 0.5 the model favors precision (0.903); use t=0.36 for higher recall at lower precision, or calibrate per use case.
Training Data
Trained on PhishTrap v2 (build 20260731T085613Z): 19,954 URLs, balanced 50/50, 70/15/15 split. Phishing URLs from PyFunceble-verified feeds (Phishing.Database, OpenPhish, PhishStats); legitimate URLs from Tranco top 10K. See the dataset card for full provenance.
Citation
@misc{saidutta69_2026_phishscout,
author = {Sai Dutta Abhishek Dash},
title = {PhishScout: Tiny On-Device Phishing URL Detector},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/models/saidutta69/PhishScout}},
note = {Trained on the auto-refreshing PhishTrap dataset}
}Built on PhishTrap + Phishing.Database, OpenPhish, PhishStats, and Tranco. MIT licensed.
