A factor-rotation rabbit hole · Sep 2026
It sounds like it should mean something: when aggressive, high-beta stocks are outrunning defensive, low-volatility ones, that’s supposed to be a “risk-on” tape. I tested it three ways — on the two ETFs built for exactly this trade, rebuilt from the stock market back to 1994, and then, because a real pattern and a real trade are two different things, by actually trying to trade it.
The reasoning is standard-issue finance, not a hot take. Frazzini and Pedersen’s “Betting Against Beta” and the whole low-volatility-anomaly literature are built on exactly this beta spread, and S&P runs two investable indexes on it — High Beta and Low Volatility slices of the S&P 500 — each with an ETF (SPHB, SPLV) tracking it since May 2011. When SPHB is beating SPLV, the story goes, risk appetite is up; when SPLV leads, the market is getting defensive.
Fifteen years of ETF data is a thin sample for a question about crashes — only two real ones fall inside it, 2020 and 2022. So I rebuilt the same spread from the underlying stocks, survivorship-free, back to 1994: every month, rank the S&P 500 by trailing beta and by trailing volatility, and hold the top-beta decile against the bottom-vol decile. That puts three more bear markets in the sample — 2000–02 and 2008 — instead of two.
Sharpe edge, 32 years
SPY’s realized Sharpe on High-Beta-leading days (0.81) vs Low-Vol-leading days (0.50). High-Beta leads at every lookback tested, 21–126 days.
2008, split by regime
SPY’s return on Low-Vol-leading days (82% of the year) vs −14.2% on High-Beta-leading days. 2000–02 and 2022 show the same split.
When I tried to trade it
Overfitting score from RealTest’s native CSCV/PBO test on the timing rule — a coin flip. Killed before it ever touched the locked holdout.
The regime definition is a straight relative-strength trend: divide one basket’s equity by the other’s and ask whether the ratio is above or below its own trailing average. “High-Beta leading” = the ratio is above its N-day average; “Low-Vol leading” = at or below it. For the 32-year version, the two baskets are synthetic: each month, rank all point-in-time S&P 500 members by trailing 252-day beta (the standard market-model beta, correlation times the ratio of volatilities) for the high-beta decile, and by trailing 252-day volatility for the low-vol decile — bottom decile, lowest vol. Equal-weight, about 50 names each. Two mutually-exclusive long-SPY sleeves, each active only on the days its own regime is on, show what SPY actually did while each side was leading — contemporaneous, not a forecast.
The clean way to check whether that’s a real structural difference or a fluke of one lookback window is to sweep the trend window and see if the gap holds up everywhere, not just at the value I happened to pick first.
SPY’s realized Sharpe while each regime is on · synthetic S&P 500 deciles, 1994–2026
Along the way, the naive high-beta decile did something extraordinary: equal-weighted, monthly-rebalanced, no cap-weighting to dilute the biggest movers, it fell 97.8% from its own peak in October 2002 — a real, mechanically legitimate result, not a data error. The stocks that measure as “highest beta” in a bear market are the ones that already crashed hardest, and an equal-weight decile keeps buying them every month on the way down. It’s the textbook betting-against-beta result: naive high-beta portfolios crash hardest exactly when beta matters most.
Synthetic high-beta decile vs low-vol decile · 1994–2026
Drawdown from prior peak · 1994–2026
Split by year, the pattern is unmistakable in the crashes and absent everywhere else.
| Year | SPY, full year | High-Beta-leading days | Low-Vol-leading days |
|---|---|---|---|
| 2000 | −9.7% | −7.8% | −5.3% |
| 2001 | −11.8% | −2.5% | −9.0% |
| 2002 | −21.6% | −5.7% | −17.1% |
| 2008 | −36.8% | −14.2% | −27.8% |
| 2022 | −18.2% | −13.7% | −2.2% |
2000–01 read oddly because Low-Vol-leading days dominated the year (72–75% of trading days) while the market bled slowly throughout; by 2002 and 2008 the gap is unambiguous. 2022 flips the label assignment — a rate-driven growth/high-beta selloff with defensives holding up — but the same mechanism: whichever side isn’t crashing gets to “lead.”
SPY’s return, split by regime · bear-market years, 1994–2026
Outside the crashes there’s no consistent edge, and the signal visibly lags sharp turns. In the 2009 V-recovery, Low-Vol-leading days earned +26.2% against High-Beta’s +0.4% — the exact opposite of what “risk appetite” framing would predict, because the early rally days were still labeled Low-Vol from the crash momentum. COVID flips it the intuitive way: Low-Vol-leading days were the crash itself (−1.3%), High-Beta-leading days were the recovery (+6.5%). Both readings are correct for their moment; neither would have told you what to do in real time.
A real contemporaneous pattern isn’t the same thing as a tradable rule, so I built the one rule the finding actually implies — long SPY only while High-Beta leads, cash otherwise — and ran it through a locked-holdout validation gauntlet rather than trusting the backtest. Split the 32-year sample into a development window (1994–2016, all tuning happens here) and a locked holdout (2017–2026, touched exactly once, if at all).
On the development window, at the lookback I’d have picked going in (63 trading days), the timing rule ties buy-and-hold on risk-adjusted return instead of beating it: 5.3% a year against 9.25% for SPY, but with less than half the drawdown (−32% vs −55%). Same MAR (0.17 both), nearly the same Sharpe (0.54 vs 0.56). It cuts return and risk in roughly the same proportion — a wash dressed up as lower volatility, not an edge.
MAR (return ÷ worst drawdown), development window 1994–2016
RealTest has a native overfitting test built for exactly this doubt: split the 22-year development window into 16 time blocks, form every way to pick 8 of them as “in-sample” and the other 8 as “out-of-sample” — 12,870 combinations — and ask how often the lookback that looks best in-sample also ranks well out-of-sample. A real, structural edge should win consistently across splits. Noise wins about half the time, by construction.
The result: an Overfitting Score of 51.4%. Coin-flip, and if anything a hair worse than 50% — with 12,870 splits the sampling noise on that number is tiny, so the in-sample-best lookback was actively uninformative about the out-of-sample rank, not just unhelpful. The reality haircut on top: in-sample Sharpes across splits ran 0.55–0.87, and the expected out-of-sample shrinkage was 0.21 — enough to erase most of whatever edge the in-sample number showed.
That result ends the test. The locked 2017–2026 holdout was never touched — zero peeks — because spending it on a rule that already failed the overfitting check would just contaminate data I might want for something that actually survives one day.
A real pattern in 32 years of data, and a coin flip the moment it has to pick a parameter to trade on. Both things are true at the same time.
Real crash signature, no tradable edge.The mechanism is real and it replicates: High-Beta-leading days are a genuinely smoother, higher-Sharpe regime for SPY, holding at every lookback from 21 to 126 days across 32 years, and every major bear market in the sample — 2000–02, 2008, 2022 — shows the same signature of Low-Vol dominating the calendar while the market bleeds. What it isn’t is a standalone timing signal: the one tradable rule it implies ties buy-and-hold rather than beating it, the parameter choice behind that rule is statistically indistinguishable from noise, and it never earned the right to spend the locked holdout that would have been the real test.
For what it’s worth, the indicator is currently reading High-Beta leading, about 10% above its own trend as of the last close in this data — a description of where the ratio sits, not a forecast, and not something to act on by itself.
InSPX / Constituency: $SPX. Monthly rank by trailing 252-day beta (Correl(ri,rm,252)×StdDev(ri,252)/StdDev(rm,252)) for the high-beta decile and trailing 252-day volatility for the low-vol decile (bottom decile). Equal-weight, ~50 names each, ~10% of the eligible universe.None of this makes the mechanism wrong. Betting-against-beta is one of the better-documented anomalies in finance, and low-volatility stocks really do end up “leading” when the tape is breaking, in a way that shows up cleanly across four separate bear markets and 32 years of data. It’s a legitimate thing to keep an eye on. What it isn’t — at least not in the simple long/cash form I could think to test — is something you can point a strategy at.