Tanushri205/llm-serving-autoscaler-enviornment-server
Sentinel-SOC is an AI-powered Security Operations Center where LLM agents act as autonomous security analysts — investigating, reasoning through adversarial noise, and mitigating real-world cyber threats using structured forensic tools.
Unlike traditional benchmarks that only evaluate whether an agent reaches the correct answer, Sentinel-SOC evaluates how the agent reasons: every tool call, every decision, and every mistake is tracked, scored, and explainable.
The environment follows current OpenEnv client/server conventions:
SentinelSOCEnvis the server-side environmentSentinelSOCClientis the HTTP client for remote usageinference.pyis the standard OpenEnv baseline inference runner- Scoring follows the Cyber Kill Chain methodology (Reconnaissance → Identification → Containment → Remediation)
🧭 Sentinel-SOC: Professional UI Overhaul
A transformation from a standard dashboard into a high-fidelity Security Operations Center (SOC) console designed for clarity, realism, and rapid decision-making.
🎨 Phase 1: SOC Design System
Implemented a cohesive dark-theme design inspired by real SOC tools:
- GitHub-dark base palette for low-eye-strain operation
- Neon accents:
- 🟢 Green → safe / resolved
- 🔴 Red → threat / critical
- 🟡 Yellow → pending / investigation
- Card-based layout for clean visual separation and readability
Outcome: A professional, cybersecurity-grade interface instead of a generic UI.
🧭 Phase 2: Structural HUD (Header + Quick Stats)
Redesigned layout to prioritize instant understanding:
- Persistent SOC Header Bar:
- System Operational
- Analyst Active
- Kill Chain Enabled
- Quick Stats Row:
- Steps Used
- Threat Status
- Investigation Phase
- Score
Outcome: Judges can understand system state in under 5 seconds.
🔍 Phase 3: Forensic Report Refinement
Enhanced the Incident Report for clarity and decision support:
- Status indicators:
- 🔴 Threat Ongoing
- 🟢 Mitigated
- Dynamic investigation fields:
- ⚠ Pending → when not yet identified
- 🟢 Resolved → when confirmed
- Context-aware confidence display
- Structured step tracking (e.g., 2 / 10)
Outcome: The system communicates investigation progress, not just data.
🧠 Final Result
Sentinel-SOC now delivers:
- Immediate clarity (quick stats + header)
- Deep reasoning visibility (timeline + reasoning tabs)
- Real-world feel (SOC-grade UI + workflows)
🎯 Positioning Statement
Sentinel-SOC is not just an interface — it is a professional environment for evaluating how AI agents investigate and respond to security incidents.
This makes Sentinel-SOC not only a benchmark, but a debugging tool for agent reasoning failures.
Current Architecture
Main modules:
- `environment.py`: procedural scenario engine, kill chain enforcement, and grading logic
- `models.py`: typed
IncidentActionandIncidentObsPydantic models - `client.py`: synchronous HTTP client for remote usage
- `inference.py`: OpenEnv-compliant baseline inference script (
[START]/[STEP]/[END]) - `server/app.py`: FastAPI server exposing
/reset,/step,/state,/grade,/history - `server/gradio_ui.py`: Professional Gradio forensic dashboard
- `tests/test_environment.py`: 9-test validation suite
Rewards
Rewards are structured to enforce kill chain methodology. Skipping phases or acting on decoy data is penalized.
Final score is clipped to [0.01, 0.99] for OpenEnv boundary compliance.
Quick Start
Remote (HTTP API)
import httpx
SERVER = "https://tanushri205-llm-serving-autoscaler-enviornment-server.hf.space"
# Reset the environment
obs = httpx.post(f"{SERVER}/reset", params={"task": "easy"}).json()
print(obs["incident_thread"]) # → live incident alert with situational report
print(obs["logs"]) # → noisy telemetry with mixed signals
# Reconnaissance — query logs
resp = httpx.post(f"{SERVER}/step", json={
"reasoning": "Ingesting all telemetry to identify credential anomalies",
"tool": "query_logs",
"parameters": "all"
}).json()
print(resp["reward"]) # → 0.10
# Identification — extract the IOC
resp = httpx.post(f"{SERVER}/step", json={
"reasoning": "sk_live_ prefix indicates production credential leak",
"tool": "extract_ioc",
"parameters": "sk_live_51M0xABcdEF67"
}).json()
print(resp["reward"]) # → 0.30
# Containment — inspect the source file
resp = httpx.post(f"{SERVER}/step", json={
"reasoning": "Credential observed in app.log — inspecting root cause",
"tool": "inspect_file",
"parameters": "app.log"
}).json()
print(resp["reward"]) # → 0.20
# Remediation — apply fix
resp = httpx.post(f"{SERVER}/step", json={
"reasoning": "Rotating and masking all exposed production credentials",
"tool": "apply_fix",
"parameters": "remediate"
}).json()
print(resp["reward"]) # → 0.40
print(resp["done"]) # → True
# Final grade
score = httpx.post(f"{SERVER}/grade").json()
print(score) # → {"score": 0.93}Local Usage (Direct Python)
from environment import SentinelSOCEnv
from models import IncidentAction
env = SentinelSOCEnv()
obs = env.reset(task="easy")
action = IncidentAction(
reasoning="Ingesting raw telemetry for initial reconnaissance",
tool="query_logs",
parameters="all"
)
obs, reward, done, info = env.step(action)
print(reward) # → 0.10
print(info["tool_result"]) # → "Log analysis complete. Credential pattern observed in..."
score = env.grade()
print(score) # → 0.0 – 1.0Baseline Inference
export HF_TOKEN="hf_..."
export API_BASE_URL="https://router.huggingface.co/v1"
export MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"
python inference.pyOutput follows the mandatory OpenEnv log format:
[START] task=easy env=sentinel-soc model=Qwen/Qwen2.5-72B-Instruct
[STEP] step=1 action=query_logs(all) reward=0.10 done=false error=null
[STEP] step=2 action=extract_ioc(sk_live_...) reward=0.30 done=false error=null
[STEP] step=3 action=inspect_file(app.log) reward=0.20 done=false error=null
[STEP] step=4 action=apply_fix(remediate) reward=0.40 done=true error=null
[END] success=true steps=4 score=0.930 rewards=0.10,0.30,0.20,0.40Tasks & Scenarios
Every reset() generates a unique episode — IPs, API keys, domains, filenames, and log timestamps are seeded fresh each time.
Hard tier introduces multiple suspicious—but incorrect—domains and IPs alongside the real C2 callback. The agent must reason through a multi-step deduction chain:
- Identify suspicious outbound network traffic in noisy telemetry
- Inspect obfuscated vendor source code containing a Base64-encoded payload
- Decode the payload to reveal the embedded C2 domain
- Correlate the decoded domain against observed UNAUTHORIZED network connections
- Distinguish the true C2 from decoy domains that appear in both logs and code
Single-step heuristics and pattern-matching fail in this scenario. The agent must perform genuine multi-source deductive reasoning.
🔍 IOC Extraction Design
Sentinel-SOC evaluates whether an AI agent can identify the correct Indicator of Compromise (IOC) among realistic noise and decoys. Each task introduces ambiguity that requires contextual reasoning — not pattern matching.
🟢 Easy — Credential Leak
Logs contain both a real production key and a decoy test key:
sk_live_... ← real credential (production exposure)
sk_test_... ← decoy (development key, harmless)The agent must distinguish production vs test credentials and select only the sensitive one.
🟡 Medium — SQL Injection
Multiple external IPs appear in access logs — only one is associated with a SQL injection payload:
94.103.42.17 → GET /user?id=1' UNION SELECT password_hash FROM credentials -- ← target
45.227.18.203 → POST /api/login (401 failed auth — suspicious but not the source)The agent must analyze query behavior and correlate attack patterns with the correct attacker IP.
🔴 Hard — Backdoor Detection
The malicious C2 domain appears in two places simultaneously:
# In network logs:
NETWORK: 172.16.x.x → malware-cdn.xyz:443 (UNAUTHORIZED) ← target
NETWORK: 172.16.x.x → attacker-domain.io:8080 (BLOCKED) ← decoy
# In source code (Base64-obfuscated):
beacon_url = base64.b64decode("bWFsd2FyZS1jZG4ueHl6").decode() ← encoded targetThe agent must decode the obfuscated payload, correlate it with network activity, and distinguish the true C2 endpoint from decoys that appear in both logs and source code.
🧠 Why This Is Challenging
- Multiple valid-looking candidates appear (decoys intentionally resemble real IOCs)
- Correct identification requires cross-source reasoning (logs + code)
- Cannot be solved by single-step extraction or pattern matching alone
🎯 Evaluation Objective
The agent is not rewarded for extraction alone, but for selecting the correct IOC based on contextual understanding within a structured investigation workflow. Wrong selections are penalized.
Actions and Observations
IncidentAction
reasoning: str # analyst chain-of-thought
tool: str # one of: query_logs, extract_ioc, inspect_file, apply_fix
parameters: str # target value (filename, IOC string, etc.)IncidentObs
logs: str # raw system/access/network telemetry
code_snippet: str # source code of relevant file
incident_thread: str # incident alert + situational phase report
status: str # current investigation status
steps_remaining: int # budget remaining
reward_signal: float # cumulative reward so farWhat Works Today
- ✅ Procedural scenario generation (non-memorizable, seeded per episode)
- ✅ Cyber Kill Chain enforcement with per-phase rewards and penalties
- ✅ Adversarial noise scaling (10% / 30% / 50% by difficulty)
- ✅ Hard task decoy ambiguity (multiple false-positive IOC candidates)
- ✅ Neutral situational awareness — no hand-holding guidance to agents
- ✅ Professional Gradio dashboard with 8-tab forensic analyst UI
- ✅ Investigation Timeline, Agent Reasoning, and Evaluation panels
- ✅ Plain-English "Explain Simply" mode for accessibility
- ✅ OpenEnv-compliant
[START]/[STEP]/[END]logging - ✅ Scores clipped to
[0.01, 0.99]for boundary compliance
Baseline Scores
Evaluation via inference.py with Qwen/Qwen2.5-72B-Instruct at temperature=0.0:
Test Suite
python tests/test_environment.py
# ✔ test_reset_produces_unique_scenarios PASSED
# ✔ test_kill_chain_ordering PASSED
# ✔ test_easy_full_solve PASSED (score: 0.93)
# ✔ test_decoy_penalty PASSED
# ✔ test_noise_scaling PASSED (easy: 6 lines, hard: 8 lines)
# ✔ test_grade_boundaries PASSED (min: 0.01, max: 0.93)
# ✔ test_all_tasks [easy] PASSED (score: 0.93)
# ✔ test_all_tasks [medium] PASSED (score: 0.95)
# ✔ test_all_tasks [hard] PASSED (score: 0.96)
# 🏆 ALL TESTS PASSEDLocal Setup
# Clone and install
git clone https://github.com/tanushri205/sentinel-soc.git
cd sentinel-soc
pip install -r requirements.txt
# Run the server
python -m uvicorn server.app:app --host 0.0.0.0 --port 7860
# Or run tests directly
python tests/test_environment.pyDocker
docker build -t sentinel-soc:latest .
docker run -p 7860:7860 \
-e HF_TOKEN="hf_..." \
-e API_BASE_URL="https://router.huggingface.co/v1" \
-e MODEL_NAME="Qwen/Qwen2.5-72B-Instruct" \
sentinel-soc:latestEnvironment Variables
HTTP Endpoints
Sentinel-SOC demonstrates how LLMs can transition from passive assistants to active, auditable security analysts — reducing mean time to detection, enforcing structured investigation methodology, and making AI-driven incident response transparent and trustworthy.
Developed for the Meta × HuggingFace × Scaler OpenEnv Hackathon 2026.
