AxeH666/meta_RL_mod
ModGuard-RL
ModGuard-RL is an OpenEnv environment for trust-and-safety triage. An agent reviews flagged content and must choose one of four actions: approve, remove, escalate, or legal_hold. Unlike a one-shot classifier benchmark, this environment models operational moderation flow: scarce escalation capacity, contradictory AI and human signals, legal-liability edge cases, and post-decision audits that can overturn an apparently reasonable early choice.
Why moderation RL matters
Real moderation systems are full of sequential tradeoffs. Approving harmful content increases downstream harm. Removing benign content suppresses speech and damages user trust. Escalation queues are expensive. Legal-hold workflows are rare but high stakes. A useful evaluation environment should capture all of those tensions together, not flatten them into a single label prediction.
ModGuard-RL turns that workflow into a compact RL problem with rich process feedback. It rewards not only final correctness, but also calibrated caution, coherent action sequencing, and resistance to misleading confidence signals.
What is novel here
- Adversarial AI behavior: the visible AI recommendation can be paired with very high confidence even when it is wrong.
- Conflicting signal design: risk level, reviewer hint, and account history can disagree in structured ways.
- Noisy reviewer hints: hints may be correct, noisy, absent, or actively adversarial.
- Operational edge cases: some episodes start with zero escalation budget, repeated escalation chains are penalized, and some legal-hold cases only become obvious after an audit.
- Post-decision audit stage: episodes can extend to 4 decisions, forcing the agent to revise an earlier choice when new signals surface.
- Anti-reward-hacking design: hidden ground truth is not exposed through the public
state()API.
Environment structure
Stages:
initial_review
├─ approve/remove -> terminal OR post_decision_audit
├─ escalate -> escalation_review (if budget remains)
├─ escalate -> post_decision_audit with penalty (if budget is already zero)
└─ legal_hold -> legal_review
escalation_review
├─ approve/remove -> terminal OR post_decision_audit
├─ escalate -> post_decision_audit with repeated-escalation penalty
└─ legal_hold -> legal_review
legal_review
├─ approve/remove/legal_hold -> terminal OR post_decision_audit
└─ escalate -> post_decision_audit with penalty
post_decision_audit
└─ any action -> terminalEpisode length is now 1 to 4 steps.
Observation and state
Core observation fields remain compact and OpenEnv-friendly:
content_categoryrisk_levelplatform_contextai_confidence_scorehuman_reviewer_hintqueue_pressurereviewer_overturn_ratestep_numbercase_historystage
Additional operational info is carried in observation.metadata, including:
ai_recommendationsignal_conflict_scoreuncertainty_indexscenario_tagsaudit_reasonescalation_budget_remainingreward_breakdownon terminal steps
Public state is intentionally operational, not answer-revealing. It includes budget, step count, repeated escalation count, audit flags, proposed resolution, and action history, but not hidden ground truth.
Reward design
The reward is continuous, clamped to [0, 1], and intentionally decomposed into interpretable parts:
reward =
correctness * 0.36 +
process * 0.18 +
hint_calibration * 0.10 +
speed * 0.10 +
consistency * 0.14 +
uncertainty_awareness* 0.12 -
overconfidence_penalty*0.12Component intuition:
correctness: final action quality relative to the hidden label.process: respects escalation budget, avoids path penalties, and matches the episode’s natural trajectory length.hint_calibration: rewards using good hints and ignoring bad ones.speed: prefers resolving easy cases quickly.consistency: rewards coherent action sequences and penalizes escalation chains.uncertainty_awareness: rewards caution when signals conflict and decisiveness when they do not.overconfidence_penalty: punishes blind trust in high-confidence AI signals when the case is adversarial.
This design makes reward hacking harder: repeated escalation and premature high-confidence decisions do not dominate the score, and public state does not leak the hidden label.
Example trajectories
1. Routine approve
step 1: initial_review
signals: low risk, aligned hint, low conflict
action: approve
result: terminal, high reward2. Budget trap with audit recovery
step 1: initial_review, zero escalation budget, conflicting signals
action: escalate
result: budget violation, forced audit path
step 2: post_decision_audit
new signal: high overturn risk, remove hint
action: remove
result: terminal, partial reward but process penalty remains3. Delayed legal requirement
step 1: initial_review
signals: high risk, misleading AI says remove with high confidence
action: escalate
step 2: escalation_review
hint: remove
action: legal_hold
step 3: legal_review
audit required: possible legal retention
action: legal_hold
step 4: post_decision_audit
final action: legal_hold
result: terminal, strong reward for uncertainty-aware correctionWhy this is challenging for LLM agents
- The highest-confidence AI signal can still be wrong.
- Reviewer hints are not uniformly trustworthy.
- Fast resolution helps on easy cases but hurts on ambiguous ones.
- Escalation is useful but limited, and repeated escalation is explicitly punished.
- The best policy depends on trajectory logic, not just the terminal label.
- Audit stages can reward changing your mind when new evidence appears.
This makes the environment a better stress test for sequential decision quality than a simple moderation classifier wrapper.
Task suite
Three explicit benchmark task wrappers are used in inference.py:
task_1_routine_triage: easy distribution, mostly short episodes, clean routine moderation.task_2_escalation_budgeting: medium difficulty with more zero-budget and signal-conflict cases.task_3_legal_liability_path: hard distribution with high-risk legal-hold pressure and delayed legal escalation cases.
Each task runs deterministic seeded episodes with different seed schedules, producing reproducible but diverse score distributions.
HF Spaces deployment
The container is server-first by default:
RUN_MODE=servestartsuvicornon0.0.0.0:${PORT:-7860}RUN_MODE=evalrunspython3 inference.py/healthreturns 200 quickly for readiness checks/reset,/step,/state,/schema,/metadata, and/wsare available
inference.py never spawns a subprocess server. If no server is reachable in evaluation mode, it falls back to an in-process environment.
Quickstart
The hackathon runner expects the image name `modguard-rl` (same as openenv.yaml and pyproject.toml). Build from the repository root:
docker build -t modguard-rl .
# or: make docker-build
# or: docker compose build # tags modguard-rl:latest per docker-compose.ymlRun the server:
docker run -p 7860:7860 modguard-rlEvaluation mode inside Docker:
docker run --rm -e RUN_MODE=eval modguard-rlLocal evaluation against a running server:
RUN_MODE=eval python3 inference.pyForced in-process evaluation:
RUN_MODE=eval FORCE_INPROCESS=1 python3 inference.pyPhase 2 preflight (before git / Hugging Face push)
Run the same checks the hackathon uses for Phase 2 (baseline inference, reward variance, openenv validate, Docker modguard-rl + RUN_MODE=eval):
uv run python tests/phase2_check.pyFaster local run (skip Docker, fewer episodes): PHASE2_QUICK=1 PHASE2_SKIP_DOCKER=1 uv run python tests/phase2_check.py
API summary
GET /health: liveness and readinessPOST /reset: starts or resets an HTTP sessionPOST /step: advances the active HTTP sessionGET /state: returns current public state for the active sessionGET /schema: returns action, observation, and state schemasGET /metadata: returns environment metadata and README contentWS /ws: OpenEnv-compatible persistent session channel
Judge-facing summary
ModGuard-RL aims to be strong for both automated validation and human review:
- strict OpenEnv schema endpoints
- HF Spaces-ready server startup
- sessionful HTTP and websocket support
- reproducible multi-seed evaluation
- adversarial and edge-case-heavy trajectories
- reward design that encourages calibrated moderation behavior instead of shortcut policies
