madhuria/patch2prod-arena
Patch2Prod Arena
Green CI is not the same as safe to ship.
Patch2Prod Arena is an OpenEnv-style environment for training agents to answer the harder production question after CI goes green:
Should this change actually ship?
Most coding agents optimize for a narrow loop: read error -> patch code -> rerun tests -> stop at green. That is useful, but not enough for production release safety.
Patch2Prod Arena trains agents to go beyond local repair. The agent is rewarded for diagnosing causal changes, reasoning about downstream blast radius, validating impacted services, and making a safe release decision.
Judge Checklist (All Linked Here)
- Environment pushed to Hugging Face Space (discoverable + runnable): Patch2Prod Arena Space
- Working training scripts (SFT + GRPO):
- training/train_sft.py
- training/train_grpo.py
- HF job launcher: scripts/launch_hf_grpo.sh
- RL framework used: Hugging Face TRL GRPO (with LoRA adapters)
- Re-runnable notebook (starter):
- Local: notebooks/sft-kaggle.ipynb
- Kaggle: SFT Kaggle Notebook
- Evidence of real training (loss/reward/grad/entropy plots):
- artifacts/grpo_log_analysis/loss.png
- artifacts/grpo_log_analysis/reward.png
- artifacts/grpo_log_analysis/grad_norm.png
- artifacts/grpo_log_analysis/entropy.png
- artifacts/grpo_log_analysis/clipped_ratio.png
- Short writeup / mini-blog:
- blog.md
- Evaluation scripts:
- training/evaluate_sft_policy.py
- training/evaluate_grpo_policy.py
What Patch2Prod Arena Tests
This is a structured agent environment with state, actions, rewards, and verifiable outcomes. The agent must:
- Diagnose a failing CI pipeline.
- Identify the causal change.
- Apply a minimal patch.
- Compute downstream blast radius.
- Run targeted validations.
- Make a release decision:
ship,block,canary,rollback, orrequest_owner_approval.
Example Scenarios
1) Auth SDK migration (contract break hidden behind green CI)
- Local failure:
authsdk.helpersno longer exportsbuild_retry_policy. - Naive fix: replace call, rerun unit tests, ship.
- Safe behavior: inspect dependency graph and contract-test downstream consumers.
- Ground truth:
mobile-gatewaycontract breaks after token-expiry format shift. - Correct release decision:
block, notifymobile-platform.
2) Payment schema compatibility
- Local failure: consumers expect
payment_status, service emits onlystatus. - Safe patch:
return {"status": p.status, "payment_status": p.status, "id": p.id}- Then rerun local tests and targeted downstream checks.
- If contracts pass, correct decision is
ship.
Environment API
POST /resetPOST /stepGET /state
Each action is one JSON object:
{
"action_type": "view_log",
"params": {"job_name": "unit-tests"}
}Available Actions
view_log(job_name)view_commit_history()view_diff(commit_id)cat(file_path)view_migration_guide(package)view_security_advisory(package)replace(file_path, search, replace)run_unit_tests(service)view_dependency_graph(service)mark_impacted_service(service, reason)run_contract_tests(service)view_ownership_map()submit_causal_change(commit, summary)submit_blast_radius(impacted_services)submit_release_decision(decision, reason, owner_to_notify)view_reward()
Reward Design
The reward is compositional rather than sparse pass/fail, so the model gets intermediate feedback through the investigation loop. It includes:
- CI repair
- Causal diagnosis
- Minimal repair quality
- Blast-radius reasoning
- Targeted downstream validation
- Release decision correctness
- Owner escalation
- Safety penalties for unsafe or unsupported behavior
It also applies per-step cost and timeout penalties to discourage long, low-signal trajectories.
Baseline vs Improved Policy
Baseline behavior:
view_log -> patch -> run_unit_tests -> shipImproved reference behavior:
view_log -> view_commit_history -> view_diff -> cat -> view_migration_guide
-> replace -> run_unit_tests -> view_dependency_graph -> submit_blast_radius
-> run_contract_tests -> submit_release_decisionThis makes the capability gap explicit: local CI repair versus evidence-based release safety.
How Training Works
Why single-step supervision
Instead of predicting a full plan in one shot, the training loop uses:
current environment state -> one JSON actionThis matches real interaction:
observe -> act -> reward -> observe -> actStage 1: SFT (action language acquisition)
SFT teaches the model to emit executable environment actions:
- JSON-only output
- valid
action_type - required
params - no markdown/prose/fences/placeholders
- rough investigation ordering
This stage converts a generic assistant into a tool-using release agent that can stay in-protocol.
Stage 2: GRPO (action selection optimization)
GRPO optimizes which action to choose in each state, not just format correctness. Candidate actions are scored by:
- JSON validity and schema validity
- matching reference action and params
- sequencing safety (e.g., avoid contract tests before local CI repair)
- clean termination immediately after JSON
In short:
- SFT answers: "Can I speak the environment action language?"
- GRPO answers: "Can I choose better actions under reward?"
Key GRPO lessons so far
Early runs exposed a major practical issue for small models: termination control.
When completions always hit max length (clipped_ratio=1, mean_terminated_length=0), reward quality collapses and updates become weak/unstable. For this task, a correct action is short structured JSON, so "valid action and stop" is part of core capability, not formatting polish.
Current mitigations include:
- stricter JSON-at-start parsing
- explicit penalties for trailing text after JSON
- shorter generation budgets
- multi-stop token handling to improve early termination
- prompt-format alignment with instruct tuning
Key Plots
Reward curve:

GRPO loss curve (real run):

GRPO reward curve (real run):

GRPO gradient norm:

GRPO entropy:

Repo Structure
- patch2prod/env.py: core environment and reward logic
- patch2prod/server.py: FastAPI server for OpenEnv-style interaction
- patch2prod/tasks.py: benchmark tasks
- training/train_sft.py: supervised fine-tuning starter
- training/train_grpo.py: GRPO / RL training starter
- training/evaluate_sft_policy.py: SFT policy evaluation (shared rollout logic)
- training/evaluate_grpo_policy.py: GRPO-labeled evaluation entrypoint
- training/generate_grpo_data.py: state-level GRPO dataset generation from env rollouts
- training/evaluate.py: scripted evaluation and trace generation
- inference.py: root-level inference script
Run Locally
Docker (recommended):
make dockerOpens at http://localhost:7860.
Local dev (API + static UI, auto-finds free ports):
pip install -e .
make devGenerate demo artifacts (non-interactive):
make demoOn Hugging Face Spaces the container binds to PORT (default 7860).
Training
SFT:
python training/train_sft.py \
--model Qwen/Qwen2.5-0.5B-Instruct \
--train data/sft_traces.jsonl \
--out outputs/sft_patch2prodGRPO (direct):
python training/train_grpo.py \
--model madhuria/patch2prod-sft-agent \
--base_model Qwen/Qwen2.5-0.5B-Instruct \
--train data/grpo_train_states.jsonl \
--out outputs/grpo_patch2prod_loraGRPO (HF job flow used in this repo):
bash scripts/launch_hf_grpo.sh l40sx1Evaluation:
python training/evaluate_sft_policy.py \
--policy baseline \
--tasks data/eval_tasks.jsonl \
--out artifacts/traces/baseline_trace.json
python training/evaluate_grpo_policy.py \
--model outputs/grpo_patch2prod_lora \
--base_model Qwen/Qwen2.5-0.5B-Instruct \
--tasks data/eval_tasks.jsonl \
--out artifacts/traces/grpo_trace.jsonDemo Flow
Recommended demo sequence:
- Pick a scenario.
- Run baseline policy.
- Run trained/reference policy.
- Compare timelines and action traces.
- Inspect blast-radius evidence.
- Inspect final release decision.
For the auth-sdk task, the key contrast is:
- Baseline: CI passes -> ships too early -> unsafe.
- Improved: CI passes -> downstream contract fails -> block release.
Why This Matters
Patch2Prod Arena evaluates capabilities that production release agents need:
- causal diagnosis
- dependency and ownership awareness
- downstream validation
- evidence-based release decisions
The goal is to move beyond "Can I make tests pass?" toward "Can this safely go to production?"
What Comes Next
- expand benchmark coverage beyond current tasks
- add richer synthetic step-level traces
- continue improving raw JSON termination behavior
- run shorter curriculum-style GRPO phases
- report raw policy vs policy-plus-validator separately
- keep improving UI for side-by-side decision evidence
Built by Madhuria Rudra.
