hwihwalab/cartpole-v1-ppo
๐ค CartPole-v1 // Physical AI & Planetary Sim-to-Real Benchmark Suite
  [](https://huggingface.co/spaces/hwihwalab/cartpole-v1-ppo) [](https://huggingface.co/hwihwalab/cartpole-v1-ppo) [](https://gymnasium.farama.org/environments/classiccontrol/cart_pole/)     
"Can Earth-Trained Reinforcement Learning Policies Survive Extraterrestrial Gravitational Shifts?" A high-precision Physical AI & Robotics Dynamics Benchmark comparing Deep Neural PPO (Proximal Policy Optimization) against Classical Optimal LQR (Linear Quadratic Regulator) across 4 planetary gravitational regimes and dynamic physical disturbances. [ ๐ English Documentation ](README.md) | [ ๐ฐ๐ท ํ๊ตญ์ด ๋งค๋ด์ผ ](README_KR.md) | [ ๐ฎ Live Interactive Web Demo ](https://huggingface.co/spaces/hwihwalab/cartpole-v1-ppo)
[!TIP] ๐ฎ Try Live in Browser (Zero Install): ๐ Open Hugging Face Spaces Live Demo ๐ฆ Official Model Hub: ๐ค hwihwalab/cartpole-v1-ppo | ๐ GitHub Repository: Hwihwa-Lab/cartpole-v1-ppo
๐ฎ Interactive Live Demo (Hugging Face Spaces)
๐ [Launch Interactive Physical AI Lab Space](https://huggingface.co/spaces/hwihwalab/cartpole-v1-ppo)
- ๐ฑ๏ธ Interactive Mouse Disturbance (Troll the AI): Click, drag, or flick on the canvas to inject real-time physical disturbance impulses (
โก ยฑXX.X N) and watch the AI catch and rebalance the pole in real-time. - ๐ช Planetary Zero-Shot Transfer: Switch seamlessly between Moon ($1.62\,\text{m/s}^2$), Mars ($3.72\,\text{m/s}^2$), Earth ($9.81\,\text{m/s}^2$), and Jupiter ($24.79\,\text{m/s}^2$).
- ๐ Live Phase Portrait ($\theta$ vs $\dot{\theta}$): Observe real-time orbital spiral convergence to the stable origin attractor $(0, 0)$.
- โก 6-Speed Simulation Deck: From $0.25\times$ ultra-slow motion up to $5.0\times\text{ Turbo}$ and $10.0\times\text{ Max}$.
โจ๏ธ Interactive Controls & Hotkey Mapping
๐๏ธ System Architecture
flowchart TB
subgraph Client_Layer ["๐ค Physical AI & Robotics Dynamics Suite (One-Screen Golden Ratio)"]
UI_Left["Left: Controller Arena (PPO vs LQR), 4-DOF Telemetry & Speed Dropdown"]
UI_Center["Center: 60FPS Canvas, Mouse Drag Force Vector & Phase Portrait Attractor"]
UI_Right["Right: Sim-to-Real Planetary Tuner (L, M, g) & Real-time Chart.js"]
end
subgraph Core_Engine ["โก Pure JS Physics & Controller Runtime (cartpole_sim.js)"]
Physics["Variable Physics Solver (Euler Integration with Dynamic L, M, g)"]
LQR_Ctrl["Classical Optimal LQR Controller (Riccati Gain Matrix u = -K*x)"]
PPO_Ctrl["Feed-Forward MLP Policy (Tanh x 2 -> Softmax Decision)"]
PhasePlot["Phase Plane Engine (ฮธ vs ฮธฬ Orbital Spiral Trajectory)"]
WeightsJSON["Exported Neural Weights (cartpole_weights.json)"]
end
subgraph Python_Backend ["๐ Python Training & Benchmark Infrastructure"]
Trainer["PPO Policy Trainer (train.py @ 25,000 steps)"]
Benchmark["Automated 1,800-Run Benchmark Engine (benchmark_experiments.py)"]
LocalServer["Zero-Dependency Local Launcher (run.py @ Port 8000)"]
TestSuite["Automated Test Harness (test_app.py - 6 Test Cases)"]
end
subgraph Hub_Distribution ["๐ Hugging Face Universal Deployment (deploy_to_hf.py)"]
Spaces["HF Spaces (Static SDK Zero-Latency Web Benchmark)"]
Models["HF Model Hub (Weights, Benchmark JSON, Model Card)"]
end
WeightsJSON --> PPO_Ctrl
Physics --> UI_Center
PPO_Ctrl --> UI_Left
LQR_Ctrl --> UI_Left
PhasePlot --> UI_Center
Physics --> UI_Right
Trainer --> WeightsJSON
Benchmark --> Models
LocalServer --> Client_Layer
Client_Layer --> Spaces
Trainer --> Models๐ Empirical Benchmark Results (1,800 Physical Episodes)
All empirical data below were generated across 1,800 physical evaluation episodes using our automated test harness (benchmark_experiments.py).
๐ช 1. Planetary Zero-Shot Generalization (Clean Nominal Environment)
๐ช๏ธ 2. Environmental Stress & Robustness Benchmark (Earth Gravity: 9.81 m/sยฒ)
๐ฌ Key Scientific Findings & Theoretical Insights
- Robustness of Non-linear Neural Policy:
- The Trained PPO agent exhibits remarkable zero-shot transfer capabilities across extreme gravity variations ($0.17g \sim 2.53g$), maintaining a 100% success rate without retraining.
- In high gravity (Jupiter: $24.79\,\text{m/s}^2$), PPO compensates by increasing actuator switching frequency to maintain angular equilibrium within $|\theta| \le 0.53^\circ$.
- Analytical Optimal Control vs Deep RL:
- Optimal LQR provides slightly tighter nominal angular deadband control ($|\theta| \approx 0.18^\circ$), while PPO maintains greater resilience under asymmetrical lateral wind bias due to non-linear policy exploration.
- Phase Space Limit Cycle Dynamics:
- Real-time phase portrait analysis demonstrates asymptotic spiral convergence toward the origin attractor $(0, 0)$ across both LQR and PPO architectures.
๐ฌ Kinematic Motion & Dynamic Behavior Analysis (How the System Actually Moved)
Based on continuous state-space trajectory logging across the 1,800 physical evaluation runs, each experimental regime exhibited distinct physical motion signatures:
- ๐ Earth Nominal ($9.81\,\text{m/s}^2$ ยท Symmetric Micro-Chattering):
- Cart Displacement: Cart stays tightly bounded within $|x| \le 0.12\,\text{m}$ around track center.
- Actuator Dynamics: Switches between $+10\,\text{N}$ and $-10\,\text{N}$ at $\approx 14.2\,\text{Hz}$ with a balanced $50.0\%\,\text{L} / 50.0\%\,\text{R}$ duty ratio.
- Pole Motion: Maintains vertical deadband of $|\theta| \le 0.26^\circ$ without macroscopic angular oscillation.
- ๐ Moon Low Gravity ($1.62\,\text{m/s}^2$ ยท Floaty Wave Overshooting):
- Cart Displacement: Cart oscillates across wider track excursions ($|x| \approx 0.45\,\text{m} \sim 0.82\,\text{m}$).
- Dynamic Mechanism: Due to reduced restoring gravity, the discrete $\pm 10\,\text{N}$ force impulse introduces angular momentum that takes longer to dissipate, producing visible low-frequency sinusoidal wave riding before settling.
- ๐ช Jupiter Extreme Gravity ($24.79\,\text{m/s}^2$ ยท High-Frequency Hyper-Stiffness):
- Actuator Dynamics: Actuator switching frequency spikes to $>22.5\,\text{Hz}$.
- Dynamic Mechanism: Gravitational torque $\tau_g = m g l \sin\theta$ amplifies $2.53\times$, forcing the neural policy to deliver rapid-fire micro-corrections to prevent tipping beyond the irreversible divergence threshold.
- ๐จ Lateral Wind Bias ($+2.2\,\text{N}$ ยท Asymmetric Lean Counter-Steering):
- Duty Cycle Shift: Policy autonomously shifts duty ratio to $64.8\%\,\text{Left} / 35.2\%\,\text{Right}$.
- Kinematic Posture: Cart holds a steady bias position at $x \approx -0.18\,\text{m}$ with pole leaning slightly upwind to balance aerodynamic drag against gravity.
- โก External Perturbation Recovery (Two-Phase Counter-Steer & Settle):
- Phase 1 (Catch): When a $+15\,\text{N}$ impulse hits, cart rapidly accelerates in the disturbance direction to position its pivot beneath the falling center of mass.
- Phase 2 (Return): Once angular velocity $\dot{\theta} \rightarrow 0$, cart slowly glides back toward origin $x = 0.0\,\text{m}$ along a stable phase-plane spiral trajectory.
๐ Repository Structure & Manifest
โก Quick Start & Local Replication
1. Launch Standalone Desktop App
python run.py2. Re-run Automated Empirical Benchmark (1,800 Episodes)
python benchmark_experiments.py3. Run System Test Suite
python test_app.py๐ Hwihwa Robotics Ecosystem Roadmap
This project represents Foundation Stage 1 in the Hwihwa Lab Physical AI & Robotics Series:
- CartPole-v1 PPO ยท 1D Classical Inverted Pendulum Dynamics & Sim-to-Real Benchmark
- LunarLander-v3 D3QN ยท 2D Dual-Thruster Lunar Descent & Vector Dynamics
- LeRobot Push-T ยท 2D Teleoperation & Diffusion Imitation Learning
- LeRobot ALOHA Sim ยท Bimanual Robotic Manipulation & Actuator Array
- MicroDuck 14-DOF ยท 3D Bipedal Digital Twin Real-Time Flight Deck
๐ License
This project is licensed under the MIT License - see the LICENSE file for details.
Trained and deployed with [CartPole Physical AI Lab](https://huggingface.co/spaces/hwihwalab/cartpole-v1-ppo) by HWIHWA LAB.
