LSTM Forecasting & Walk-Forward Validation
View Repository SourceA 1-minute OHLCV forecasting pipeline enforcing causal EWMA standardization, leading to the rejection of a statistically significant model due to unexplainable edge.
I built a modular, 1-minute OHLCV forecasting pipeline for SPY/QQQ/VXX. A key focus was ensuring strict causal session-reset EWMA standardization. This guarantees train/serve parity by construction without the risk of retroactively fitting scalers and leaking future data.
Defensive Data Engineering
I enforced strict timezone and session contracts that hard-fail on non-contiguous days or extended-hours data. The pipeline constructed 19 causal features per step, including lead–lag correlation matrices across symbol pairs and log vol-surge interaction terms.
Rejecting the “Winner”
I executed an audited full-history walk-forward A/B test (29 folds/arm) on a remote VPS. The winning fixed-horizon model achieved a statistically significant edge ($p < 10^{-6}$).
However, I actively rejected this winning model for production default. By probing the network, I proved that its “edge” derived primarily from a volatility-nowcast shortcut rather than the intended directional signal. An edge you cannot explain is usually a shortcut; statistical significance is not economic significance.
Production Integration
Despite rejecting the primary predictive arm, the pipeline infrastructure proved sound. I deployed a live VXX/SPY signaling engine against Alpaca paper/live endpoints for momentum execution, supported by comprehensive real-data validation.