Angshuman12/Anvil
<p align="center"> <img src="https://capsule-render.vercel.app/api?type=waving&color=0:050e1a,50:0f766e,100:14b8a6&height=200§ion=header&text=ANVIL&fontSize=80&fontColor=ffffff&fontAlignY=38&animation=fadeIn&desc=Adversarial+Neural+Vulnerability+Inspection+and+Learning&descSize=14&descAlignY=58&descColor=5eead4" /> </p>
<p align="center"> <img src="https://img.shields.io/badge/Python-3.10+-3776AB?logo=python&logoColor=white" /> <img src="https://img.shields.io/badge/PyTorch-2.x-EE4C2C?logo=pytorch&logoColor=white" /> <img src="https://img.shields.io/badge/Gemini-2.5Flash-4285F4?logo=google&logoColor=white" /> <img src="https://img.shields.io/badge/LangGraph-StatefulAgent-16a34a" /> <img src="https://img.shields.io/badge/UMAP+HDBSCAN-Clustering-8b5cf6" /> <img src="https://img.shields.io/badge/FastAPI-REST_API-009688?logo=fastapi&logoColor=white" /> <img src="https://img.shields.io/badge/HuggingFace-Spaces-FFD21E?logo=huggingface&logoColor=black" /> <img src="https://img.shields.io/badge/License-MIT-22c55e" /> </p>
<h3 align="center">Autonomous ML Red-Teaming · Attack · Cluster · Explain · Patch · Report</h3>
<p align="center"> <a href="https://ganglet.github.io/Anvil"><strong>Live Demo →</strong></a> | <a href="https://angshuman12-anvil.hf.space"><strong>API →</strong></a> </p>
What it does
ANVIL takes any PyTorch neural network, runs it through a fully autonomous 8-phase adversarial auditing pipeline, and produces a professional PDF audit report — with zero human decisions.
<p align="center"> <img src="docs/architecture.png" alt="ANVIL Architecture Diagram" width="360" /> </p>
Pipeline Deep Dive
Phase 1 — Model Interface
Any PyTorch network subclasses BaseModel and exposes three methods. ResNet-18 and DistilBERT ship as first-party wrappers.
class ImageModel(BaseModel):
def predict(self, x: Tensor) -> Tensor: ... # → logits
def get_gradients(self, x, y) -> Tensor: ... # → ∂L/∂x
def get_activations(self, x) -> Tensor: ... # → penultimate layerPhase 2 — Attack Surface Profiler
Phase 3 — Attack Engine
Four attack strategies, all implemented from scratch in PyTorch autograd:
Each successful attack produces an AdversarialExample carrying the original tensor, perturbed tensor, true label, predicted label, attack name, epsilon, and per-sample confidence scores.
Phase 4 — Failure Mode Clustering
Penultimate-layer activations encode why the model was fooled — not just that it was fooled. UMAP projects these high-dimensional vectors to a low-dimensional embedding preserving local manifold structure. HDBSCAN then finds density-based clusters without a fixed cluster count.
n_neighbors = min(15, N - 1)
n_components = min(5, N - 1)
If UMAP raises scipy.linalg.eigh (N < 20) → fallback to PCA(n_components=2)
Noise points (cluster = -1) are counted but not explainedPhase 5 — LLM Explanation Agent
A stateful LangGraph agent with FAISS retrieval over 10 adversarial ML papers:
Goodfellow et al. 2015 · Madry et al. 2018 · Carlini & Wagner 2017 · Brown et al. 2017 · Szegedy et al. 2014 · Papernot et al. 2016 · Xie et al. 2019 · Cohen et al. 2019 · Zhang et al. 2019 · Croce & Hein 2020
For each cluster the agent injects cluster statistics (centroid, attack distribution, member count) + retrieved paper chunks into a structured prompt. Gemini 2.5 Flash generates a grounded explanation with recommended patch strategy. The state graph can revisit reasoning if a coherence check fails.
Phase 6 — Autonomous Patching
Safety gate formula:
$$score = 0.6 \times resistance\gain + 0.4 \times accuracy\retention$$
A patch is accepted only if score ≥ 0.70 AND accuracy drop ≤ 3%. On failure the engine escalates to the next strategy (up to 3 attempts per cluster).
Phase 7 — Audit Report
ReportLab generates a multi-page PDF: cover page with audit metadata, executive summary, matplotlib radar chart of per-attack success rates (flat polygon = uniform robustness, spiked polygon = asymmetric weakness), per-cluster cards with LLM explanations and patch outcomes, methodology appendix.
Phase 8 — REST API
POST /audit/upload multipart: files[] + model + budget → { job_id }
GET /audit/job/{id} → { status, vulnerability_score, clusters_found, report_filename }
GET /report/{filename} → PDF stream
GET /health → { status: "ok" }Async job management via FastAPI BackgroundTasks. In-memory job store with polling. CORS configured for ganglet.github.io. Deployed via Docker on HuggingFace Spaces (free tier, 16 GB RAM).
Stack
Quick Start
git clone https://github.com/Ganglet/Anvil
cd Anvil/Anvil_Project
pip install -r requirements.txt
# Run the API server
uvicorn api:app --host 0.0.0.0 --port 8000
# Or run the CLI pipeline directly
python run.py --model resnet18 --budget 50Use the hosted demo at [ganglet.github.io/Anvil](https://ganglet.github.io/Anvil) — upload images, get a full PDF audit report back.
Project Structure
Anvil_Project/
├── models/ # Phase 1 — BaseModel ABC + ResNet-18/DistilBERT wrappers
├── profiler/ # Phase 2 — AttackSurfaceProfiler (Captum)
├── attacks/ # Phase 3 — FGSM, PGD, Patch, Semantic + AttackEngine
├── clustering/ # Phase 4 — FeatureExtractor, FailureModeClusterer, VulnerabilityTaxonomy
├── agent/ # Phase 5 — LangGraph agent, FAISS RAG, Gemini integration
├── patching/ # Phase 6 — Patcher, 4 strategies, safety gate
├── reporter/ # Phase 7 — ReportLab PDF generation
├── api.py # Phase 8 — FastAPI server with async job management
├── run.py # CLI entry point
├── requirements.txt
├── Dockerfile
└── frontend/ # React + Vite source for ganglet.github.io/AnvilAcknowledgments
- PyTorch — pytorch.org
- Captum — pytorch.org/captum
- UMAP-learn — github.com/lmcinnes/umap
- HDBSCAN — github.com/scikit-learn-contrib/hdbscan
- LangGraph — langchain.com/langgraph
- LangChain — langchain.com
- FAISS — github.com/facebookresearch/faiss
- ReportLab — reportlab.com
- Matplotlib — matplotlib.org
- FastAPI — fastapi.tiangolo.com
- HuggingFace — huggingface.co
- Adversarial ML Benchmark (ATB) — github.com/MadryLab/ATB
<p align="center"> <img src="https://capsule-render.vercel.app/api?type=waving&color=0:8e44ad,50:2980b9,100:95a5a6&height=120§ion=footer" /> </p>
License
Code: MIT — see LICENSE
<p align="center"> <img src="https://capsule-render.vercel.app/api?type=waving&color=0:2980b9,50:922b21,100:c0392b&height=120§ion=footer" /> </p>
