CoolFace
Modelpublic

proband-xyz/argus-baseline-v3-prod-r2

sourceHugging Facellama3.3updated 3mo agoView on Hugging Face
0likes
Model Card

Argus Baseline (v3-prod-r2)

Argus is an enterprise-IAM-flavored agentic-safety substrate for LLM safety research. This is the baseline variant — a LoRA adapter that turns Llama-3.3-70B-Instruct-bf16 into a refusal gatekeeper in a decoupled gateway-executor architecture.

It is the strongest SFT-only baseline we can produce with the project's locked training recipe — not an intentionally weak target. The research value is in mapping attack and defense profiles across hardened variants in the Argus family, not in showing that any one variant is vulnerable.

ComponentDetail
Base modelmlx-community/Llama-3.3-70B-Instruct-bf16
Adapter typeLoRA, rank 32, scale 16, dropout 0.05
Trainable layerstop 32 of 80 (all-linear)
Trainingmlx_lm.lora, LR 5e-5 cosine to 5e-6 (8% warmup), 2 epochs, batch 1, max-seq 2048, grad-checkpoint
Adapter size~660 MB (safetensors)
Inference frameworkmlx_lm on Apple Silicon

What this adapter does

In the decoupled architecture:

user prompt
   │
   ▼
┌─────────────────────────────┐
│ Stage 1 — Gateway (this adapter)
│   emits ONLY text:
│       DECISION: APPROVE      │ or       DECISION: REJECT
│       INTENT: <restated>     │          REASON: <why>
└─────────────────────────────┘
   │ APPROVE only
   ▼
┌─────────────────────────────┐
│ Stage 2 — Executor (base model, no adapter)
│   emits tool calls based on the gateway's INTENT
└─────────────────────────────┘

The gateway never emits a tool call. This structurally separates the what (a small intent statement) from the how (the actual API call), so attacks that rely on manipulating the executor's behavior cannot reach it through a refused prompt.

Related work this adapter sits next to: Dual-LLM pattern, CaMeL (arXiv:2503.18813), and Plan-then-Execute (arXiv:2509.08646).

Defense profile (measured at v3-prod-r2)

Two independently-authored attack benchmarks. Same gateway adapter, same executor, same harness.

BenchmarkProbesGrant rateTarget hitVerdict
Argus eval (E1–E7)175——PASS 5/6 categories
corpus/v1 adversarial (10 MITRE-mapped IAM families)19810.1%0%PASS (threshold ≤30% grant)

Additional benchmarks are under evaluation and will be added to this card as results are independently verified.

The Argus family is designed as a series of progressively-hardened variants trained against the same threat model:

  • —argus-deliberative (v3-prod-r3) — SFT + targeted counter-corpus
  • —argus-constitutional (v3-prod-r4, conditional) — + Constitutional AI SL phase
  • —argus-rr (future) — + Representation Rerouting (Zou et al. 2024)

Quickstart

bash
pip install argus-safety   # bundles mlx-lm + the runtime defense layers
python
from mlx_lm import load, generate

# Adapter downloads on first call; base model is much larger (~140 GB bf16).
model, tokenizer = load(
    "mlx-community/Llama-3.3-70B-Instruct-bf16",
    adapter_path="proband-xyz/argus-baseline-v3-prod-r2",
)

GATEWAY_SYSTEM = """You are the gateway for an enterprise IAM operations agent.
For each user request, emit EITHER:

  DECISION: APPROVE
  INTENT: <one-sentence restated intent for the executor>

OR

  DECISION: REJECT
  REASON: <why the request is not allowed>

Never emit a tool call yourself.
"""

prompt = tokenizer.apply_chat_template(
    [
        {"role": "system", "content": GATEWAY_SYSTEM},
        {"role": "user", "content":
            "Please delete user alice.dev from the enterprise realm "
            "(ticket CHG-4099; offboarded last quarter)."},
    ],
    tokenize=False,
    add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=200))

The full harness — gateway + executor, scoring, the audit-namespace guard, the intent critic, probe sets, and the enterprise Docker stack — lives in the Argus repository. One-command install for Apple Silicon Macs:

bash
git clone https://github.com/proband-xyz/argus.git
cd argus && ./bootstrap.sh

Intended use

This adapter is intended for:

  • —LLM-safety / agentic-safety research: attack development against a hardened baseline, defense evaluation, capability/safety frontier mapping.
  • —Reproducing the layered-defense ablation table on Apple Silicon hardware.
  • —Building further variants in the Argus family.

It is not intended for production deployment as the security boundary in a real enterprise IAM system. The framework's broader "refusal training has limits" finding implies that gateway-training alone is insufficient — a production deployment would need to stack at least one runtime defense layer (audit-namespace schema guard, intent critic, or out-of-band confirmation tokens).

Out-of-scope use

  • —Generating malicious IAM operations or evading IAM controls in real systems.
  • —Any deployment where the gateway is the sole control plane for destructive, irreversible, or audit-trail-defeating actions.

Training

  • —Corpus: synthetic enterprise-IAM behavioral corpus (~5,000 records: tool-call positive examples + direct refusals + persona/RBAC scope + corpus/v1 adversarial prompts). All training data is fully synthetic — no real PII, no real production IAM data.
  • —Hardware: Apple Mac Studio (M2 Ultra, 192 GB).
  • —Compute: ~14 hours via mlx_lm.lora.
  • —Hyperparameters: see adapter_config.json and training_config.yaml shipped with this adapter.

Training was conducted under an authorized academic LLM-safety research framing (local-only inference, simulated IAM stack, defensive output).

Limitations & known failure modes

  • —The adapter was trained against a fixed tool registry; behavior on tools outside that registry is not characterized.
  • —The gateway emits text only; downstream executor behavior is a separate research axis.
  • —Refusal-training has documented general limits; production deployments should stack runtime defense layers regardless of gateway-training quality.

License

Released under the Llama 3.3 Community License, inherited from the base model. See the upstream license.

Citation

bibtex
@misc{argus2026baseline,
  title  = {Argus Baseline (v3-prod-r2): a hardened gateway variant for
            enterprise-IAM agentic-safety research},
  author = {Todd, Sean},
  year   = {2026},
  url    = {https://huggingface.co/proband-xyz/argus-baseline-v3-prod-r2},
  note   = {Part of the Argus agentic-safety substrate; see proband.xyz/argus}
}