madhuria/patch2prod-arena
fix: clean deployment, dev workflow, and docs for HF Space
fix(space): align docker startup with Hugging Face app port
docs: add local and Kaggle SFT notebook references
artifacts: add GRPO clipped ratio plot via LFS
artifacts: add GRPO completion plot via LFS
artifacts: add GRPO entropy plot via LFS
artifacts: add GRPO grad norm plot via LFS
artifacts: add GRPO reward plot via LFS
artifacts: add GRPO loss plot via LFS
docs: fix markdown image rendering and expand experiment log
docs: add project blog writeup for hf submission
update README and add GRPO training log
fix: gate kl_coef on GRPOConfig signature to avoid TypeError
fix: eval/train format alignment + KL + live metrics + multi-panel plots
fix: lightweight GRPO training - chat template + stop tokens + tuned hyperparams
chore: trigger hf space rebuild for README sync
Readme updates
fix hf short_description length
update hf space README metadata fields
fix hf space config: use root index.html app_file
fix hf spaces metadata color
update demo ui and add ci webhook + hf space setup
fix GRPO training: strict parse, termination reward, variance jitter, throughput
grpo training
grpo training
Fix GRPO training: use env-grounded prompts, add reward debug logging, honest eval
evaluate grpo
grpo changes
grpo changes
setting up grpo
Post SFT Training
Add working baseline and risk-aware evaluation traces
training fixes
training stuff
Add baseline evaluation artifacts
Initial Patch2Prod Arena submission
