mohareddy1423/PrivacyOps-X-final
PrivacyOps-X
PrivacyOps-X is an OpenEnv benchmark for safety-critical privacy operations. Instead of treating privacy compliance like a generic chatbot problem, this project turns it into a structured environment in which an agent must gather evidence, consult policy, coordinate with reviewers, communicate safely, and submit an auditable final resolution.
The benchmark is built around real privacy-ops failure modes:
- uncertain identity and guardian authority
- deletion requests that conflict with retention or legal hold
- fraud-review and audit-review requirements
- adversarial requester instructions that try to bypass process
- long-horizon workflows that require multiple dependent steps
The result is a benchmark that is not just interactive, but trainable, measurable, and inspectable.
Deliverables
- Hugging Face Space (submitted app): mohareddy1423-privacyops-x-final.hf.space
- Hugging Face Space repo: mohareddy1423/PrivacyOps-X-final
- GitHub repo: Mohanreddy-lab/PrivacyOps-X
- Training notebook: `notebooks/privacyops_x_trl_colab.ipynb`
- Training guide: `TRAINING.md`
- Full writeup / blog: `blog.md`
- Submission notes: `ROUND2_SUBMISSION.md`
Why this benchmark exists
Many AI demos can produce impressive language while still failing the real task. Privacy operations is exactly the kind of domain where that gap matters:
- a fluent answer can still be legally wrong
- a helpful tone can still hide an unsafe deletion
- a model can sound decisive while missing key evidence
- a one-shot response can bypass the multi-review workflow an analyst must actually follow
PrivacyOps-X was designed to measure whether an agent can behave like a careful privacy analyst under constraint, not just whether it can sound confident.
What a user does in PrivacyOps-X
In the live app, the user or evaluator interacts with a privacy case in a structured loop:
- choose a task and deterministic variant
- reset the environment
- inspect the ticket and current workspace state
- take structured actions such as opening records or searching policy
- request legal, compliance, or audit review when needed
- draft safe communication
- submit a final resolution and receive a score plus failure analysis
That means the user experience is closer to a privacy operations console than a normal chat window.
What the benchmark exposes
Action space
The acting agent uses explicit actions rather than free-form tool calls:
inspect_caseopen_recordsearch_policyopen_policy_articleset_case_fieldadd_internal_notedraft_replymessage_requesterrequest_reviewself_reviewsubmit
Observation space
Each step returns rich operational state, including:
- ticket summary
- workspace state
- visible records
- visible policy articles
- requester thread
- stakeholder inbox
- review findings
- milestones
- theme alignment
- explanation trace
- risk score
- steps remaining
- draft reply
- improvement lessons
This is what makes the environment debuggable and trainable. When the agent fails, we can inspect why it failed instead of only seeing a bad score.
Core benchmark themes
PrivacyOps-X explicitly targets the four Round 2 themes:
- Multi-agent interactions: requester, legal, compliance, audit, and critic-style feedback loops
- Long-horizon planning: milestone tracking across triage, evidence, policy, review, and resolution
- World modeling: jurisdiction, identity, legal hold, retention, fraud review, and linked-account state
- Self-improving agents: post-episode lessons and curriculum-style capability growth
Public tasks
The benchmark ships with public tasks that progressively increase operational difficulty:
- Easy: verified access request with injection-style policy bypass attempt
- Medium: unverified erasure request across multiple accounts with retention complications
- Hard: guardian/minor request with legal hold and fraud review
- Finale showcase: cross-border recovery cascade with linked records, review coordination, and partial-fulfillment logic
Key quantitative results
These numbers tell a clear story:
- random behavior is not enough
- teacher policy establishes a strong ceiling
- self-improvement produces a large and stable gain
- the 1.7B training pipeline successfully learns the benchmark distribution
Key plots
1.7B SFT loss curve
Self-improvement curve
Baseline comparison
How the system works
1. Deterministic environment loop
PrivacyOps-X uses the standard OpenEnv reset(), step(), and state() lifecycle. On reset, it instantiates a fresh privacy case with deterministic task state. On each step, it:
- validates the action
- updates hidden and visible state
- updates risk and milestone progress
- records reviewer findings and explanation trace
- computes dense reward and end-of-episode score breakdowns
2. Reviewer-backed operational logic
The benchmark is explicitly multi-agent in structure:
- Privacy analyst agent acts in the environment
- Requester agent reveals facts during follow-up
- Legal reviewer checks retention and legal-hold logic
- Compliance reviewer checks privacy workflow correctness
- Audit reviewer checks defensibility and process quality
- Critic / improvement layer emits lessons after failure
3. Measurable failure modes
Episodes track structured failure modes such as:
- hallucination
- policy violation
- logic error
- unsafe action
- redundancy
- verification error
- evidence gap
- overconfidence
- requester miscommunication
This is one of the strongest parts of the project: it does not only score outcomes, it explains what went wrong operationally.
Training pipeline
PrivacyOps-X includes a full training and evaluation workflow:
- generate teacher trajectories with
scripts/generate_sft_dataset.py - evaluate baseline policies with
scripts/evaluate_policies.py - fine-tune a policy with
scripts/train_trl_sft.py - optionally train directly against the environment with
scripts/train_openenv_grpo.py - plot results with
scripts/plot_eval_results.py - run an explicit self-improvement loop with
scripts/run_self_improvement_cycle.py
For judge-friendly reproducibility, the repo includes both:
- a script-based pipeline
- a rerunnable Colab notebook in `notebooks/privacyops_x_trl_colab.ipynb`
Recovered 1.7B training evidence
A final Colab run produced a recoverable export bundle containing:
- the LoRA adapter checkpoint
- training log history
- checkpoint snapshots at step
75and150 - tokenizer artifacts
- the loss-curve image committed in this repo as evidence
The recovered training run showed:
- smooth optimization
- sharp loss reduction
- near-saturated token accuracy
- completion of the full
150/150training schedule
This confirms that the benchmark is not only a hand-authored environment; it can also support a real post-training workflow.
Important implementation note
The original Colab post-training evaluation failed because the model emitted an older JSON action shape using keys like action, record_id, and field. The repository now includes a compatibility normalization path in `scripts/evaluate_policies.py` so recovered checkpoints can be evaluated more safely without another full retraining run.
Repository map
- `server/env.py` — core environment logic
- `server/app.py` — FastAPI app and interactive endpoints
- `server/teacher.py` — teacher trajectories and oracle policy
- `models.py` — typed action / observation / state models
- `TRAINING.md` — end-to-end training instructions
- `notebooks/privacyops_x_trl_colab.ipynb` — Colab training notebook
- `blog.md` — full project writeup
Running locally
pip install -e .[train]
python scripts/generate_sft_dataset.py --output outputs/train/privacyops_x_sft.jsonl
python scripts/evaluate_policies.py --policy random --output outputs/evals/random.json
python scripts/evaluate_policies.py --policy teacher --output outputs/evals/teacher.json
python scripts/run_self_improvement_cycle.py --task-id finale_cross_border_recovery_cascade --output outputs/evals/self_improvement_cycle.json --plot-output outputs/plots/self_improvement_curve.pngThen launch the app:
uvicorn server.app:app --host 0.0.0.0 --port 8000Honest assessment
PrivacyOps-X is strongest when it is presented honestly:
- it is a benchmark and training environment, not an autonomous production deletion bot
- its best empirical result is currently the self-improvement jump from `0.6087` to `0.9519`
- the recovered 1.7B training run is strong evidence that the SFT pipeline works
- the post-training evaluation artifact for that exact run was interrupted in Colab, but the checkpoint and training evidence were recovered
That honesty actually makes the project stronger. The core value of the work is the benchmark design, the explicit reviewer structure, the measurable improvement story, and the fact that the environment supports both evaluation and training.
Why this matters
Privacy compliance is exactly the kind of domain where AI systems need more than style:
- they need evidence
- they need constraint tracking
- they need safe communication
- they need multi-step planning
- they need auditable outputs
PrivacyOps-X demonstrates that this space can be framed as a deterministic, trainable, measurable agent benchmark.
Read next
If you want the full story, design rationale, experimental interpretation, and benchmark framing, read the full writeup here:
➡️ `blog.md`
