CoolFace
Modelpublic

graceesthi/ug-cppo-finai-2025

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes471downloads
Model Card

UG-CPPO v3: Uncertainty-Gated CVaR-PPO Trading Agents (Multi-Seed, Honest Eval)

Trained models for the UG-CPPO paper v3 (FinAI Contest 2025, NeurIPS 2026 submission).

Author — Grace-Esther Dong · Aivancity Paris-Cachan Paper — `UG_CPPO_preprint_PAT_corrected.pdf` Code — https://github.com/graceesthi/ug_cppo Preprint — arXiv [TBA]

What's in this repo

30 trained Stable-Baselines3 agents (10 seeds × 3 algorithms):

  • —Seeds: 42, 43, 44, 45, 46, 47, 48, 49, 50, 51
  • —Training: 250,000 timesteps each
  • —Evaluation: 2019-2023 Nasdaq data
  • —Tickers: 10 stocks (AAPL, MSFT, AMZN, NVDA, META, GOOGL, TSLA, NFLX, AMD, COST)

File naming

{mode}_seed{seed}.zip
  ppo_seed42.zip       → Vanilla PPO, seed 42
  cppo_seed42.zip      → CVaR-PPO, seed 42
  ug_cppo_seed42.zip   → UG-CPPO (ours), seed 42
  ... (3 modes × 10 seeds = 30 files total)

Results (250k steps, 10 seeds, honest multi-seed eval)

Cumulative Return (mean ± std)

ModelMeanStdRachevMDDWilcoxon p (vs PPO)
PPO43.94%±32.18%0.9445−27.95%—
CPPO39.71%±46.01%0.9408−31.08%0.1720
UG-CPPO35.99%±38.70%0.9420−29.72%0.8127

Interpretation:

  • —UG-CPPO cumulative return is −7.95pp lower than PPO (95% CI includes zero)
  • —Wilcoxon rank-sum test: p=0.8127 >> 0.05 → no significant difference in medians
  • —H2 hypothesis (UG-CPPO > PPO): not rejected but also not accepted (honest null-preserving stat)
  • —Honest variance (σ=38.7%) reflects genuine seed-to-seed variability

Top performers (by Rachev):

  • —Seed 47 (UG-CPPO): Rachev 1.0104
  • —Seed 46 (UG-CPPO): Rachev 0.9940
  • —Seed 51 (PPO): Rachev 0.9915

Quick load

python
from stable_baselines3 import PPO
from huggingface_hub import hf_hub_download

# Download UG-CPPO seed 47 (top performer)
path = hf_hub_download(
    repo_id="graceesthi/ug-cppo-finai-2025",
    filename="ug_cppo_seed47.zip"
)
agent = PPO.load(path)

Reproducibility

  • —Hardware: Apple M-series (CPU only)
  • —Config: 250k steps, 10 independent runs (seeds 42-51)
  • —Hyperparams: lr=1e-3, batch_size=128, γ=0.99, CVaR α=0.05
  • —Statistical test: Wilcoxon rank-sum (non-parametric, no normality assumption)

Files

  • —ppo_seed*.zip, cppo_seed*.zip, ug_cppo_seed*.zip — Trained agents
  • —multiseed_report_v13.json — Full results with Wilcoxon tests
  • —UG_CPPO_paper.pdf — Full paper with methodology
  • —multiseed_performance.png — Performance comparison plot

Citation

bibtex
@inproceedings{dong2026ugcppo,
  title={UG-CPPO: Uncertainty-Gated LLM Infusion for Risk-Sensitive
         Reinforcement Learning Trading Agents},
  author={Dong, Grace-Esther},
  booktitle={NeurIPS 2026 — FinAI Contest 2025, Task 1},
  year={2026},
  note={v3: multi-seed honest evaluation with PAT corrections}
}