PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading
Organizations: College of Engineering and Computer Science, VinUniversity Ha Noi, Viet Nam
Abstract
Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return and mean Sharpe ratio . Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.
Figures & tables
| Regime | Condition | |
|---|---|---|
| Bull, calm | bull and not stress | 1.0 |
| Bull, stressed | bull and stress | 0.5 |
| Neutral | mixed trend | 0.3 |
| Neutral-negative | weak negative trend | 0.1 |
| Bear, calm | bear without stress | 0.0 |
| Bear, stressed | bear and stress | -0.1 |
| Input: historical market data, VIX, feature set, . |
| Initialize PPO actor-critic parameters. |
| for each training timestep do |
| Observe . |
| Compute regime from trend and VIX z-score. |
| Compute target exposure . |
| Sample raw PPO action . |
| Parameter | Value |
|---|---|
| Total timesteps | 30,000 |
| Main seed | 42 |
| Learning rate | |
| Rollout steps | 1024 |
| Batch size | 128 |
| PPO epochs/update | 10 |
| Method | Formulation |
|---|---|
| Buy and Hold | after the initial allocation. |
| Risk parity | . |
| CPPI | . |
| PPO/SAC profit | . |
| PPO variance | . |
| PPO static MDD | . |
| Method | Total Return | Ann. Return | Sharpe | Sortino | Calmar | MDD |
|---|---|---|---|---|---|---|
| Buy and Hold | 0.1760 | 0.0556 | 0.3420 | 0.4251 | 0.1630 | -0.3410 |
| Risk-Parity | 0.0799 | 0.0260 | 0.2281 | 0.2848 | 0.0898 | -0.2892 |
| CPPI | -0.0275 | -0.0093 | -0.0021 | -0.0026 | -0.0427 | -0.2174 |
| PPO Profit | 0.1760 | 0.0556 | 0.3420 | 0.4251 | 0.1630 | -0.3410 |
| SAC Profit | 0.1760 | 0.0556 | 0.3420 | 0.4251 | 0.1630 | -0.3410 |
| PPO Variance | 0.1356 | 0.0434 | 0.3384 | 0.4208 | 0.1829 | -0.2371 |
| Method | Ret. Mean | Ret. Std | Sharpe Mean | Sharpe Std | Calmar Mean | Calmar Std | MDD Mean | MDD Std |
|---|---|---|---|---|---|---|---|---|
| PPO-HRAP | 0.2725 | 0.0109 | 0.6219 | 0.0565 | 0.4508 | 0.0499 | -0.1876 | 0.0210 |
| PPO Static MDD | 0.0575 | 0.2008 | 0.1320 | 0.3882 | 0.0550 | 0.1829 | -0.3554 | 0.0321 |
| PPO Profit | 0.1781 | 0.0059 | 0.3444 | 0.0067 | 0.1648 | 0.0053 | -0.3413 | 0.0005 |
| PPO Variance | -0.1126 | 0.1062 | -0.3720 | 0.2117 | -0.1512 | 0.0851 | -0.2264 | 0.1560 |
| SAC Profit | 0.1978 | 0.0897 | 0.8644 | 0.2417 | 1.4733 | 0.6717 | -0.0469 | 0.0199 |
| Asset | Method | Total Ret. | Ann. Ret. | Sharpe | Sortino | Calmar | MDD | Turnover |
|---|---|---|---|---|---|---|---|---|
| QQQ | Buy and Hold | 0.2306 | 0.0717 | 0.3826 | 0.5082 | 0.2014 | -0.3562 | 1.0000 |
| QQQ | Risk Parity | 0.1317 | 0.0422 | 0.2985 | 0.3954 | 0.1512 | -0.2788 | 3.8472 |
| QQQ | CPPI | -0.0072 | -0.0024 | 0.1010 | 0.1361 | -0.0067 | -0.3605 | 9.7168 |
| QQQ | PPO Static MDD | 0.2306 | 0.0717 | 0.3826 | 0.5082 | 0.2014 | -0.3562 | 1.0000 |
| QQQ | PPO Profit | 0.2306 | 0.0717 | 0.3826 | 0.5082 | 0.2014 | -0.3562 | 1.0000 |
| QQQ | PPO Variance | -0.1770 | -0.0630 | -0.5244 | -0.8868 | -0.1951 | -0.3227 | 35.3054 |
| Return | Sharpe | Calmar | MDD | Filter | |
|---|---|---|---|---|---|
| 0.3 | 0.0686 | 0.3770 | 0.1792 | -0.1889 | no |
| 0.4 | 0.0729 | 0.4146 | 0.1941 | -0.1854 | no |
| 0.5 | 0.0771 | 0.4537 | 0.2126 | -0.1787 | yes |
| 0.6 | 0.0800 | 0.4863 | 0.2299 | -0.1713 | yes |
| 0.7 | 0.0767 | 0.4835 | 0.2291 | -0.1650 | yes |