A factor-rotation rabbit hole · Sep 2026

Does high beta leading low vol mean anything?

It sounds like it should mean something: when aggressive, high-beta stocks are outrunning defensive, low-volatility ones, that’s supposed to be a “risk-on” tape. I tested it three ways — on the two ETFs built for exactly this trade, rebuilt from the stock market back to 1994, and then, because a real pattern and a real trade are two different things, by actually trying to trade it.

The reasoning is standard-issue finance, not a hot take. Frazzini and Pedersen’s “Betting Against Beta” and the whole low-volatility-anomaly literature are built on exactly this beta spread, and S&P runs two investable indexes on it — High Beta and Low Volatility slices of the S&P 500 — each with an ETF (SPHB, SPLV) tracking it since May 2011. When SPHB is beating SPLV, the story goes, risk appetite is up; when SPLV leads, the market is getting defensive.

Fifteen years of ETF data is a thin sample for a question about crashes — only two real ones fall inside it, 2020 and 2022. So I rebuilt the same spread from the underlying stocks, survivorship-free, back to 1994: every month, rank the S&P 500 by trailing beta and by trailing volatility, and hold the top-beta decile against the bottom-vol decile. That puts three more bear markets in the sample — 2000–02 and 2008 — instead of two.

Sharpe edge, 32 years

+0.31

SPY’s realized Sharpe on High-Beta-leading days (0.81) vs Low-Vol-leading days (0.50). High-Beta leads at every lookback tested, 21–126 days.

2008, split by regime

−27.8%

SPY’s return on Low-Vol-leading days (82% of the year) vs −14.2% on High-Beta-leading days. 2000–02 and 2022 show the same split.

When I tried to trade it

51%

Overfitting score from RealTest’s native CSCV/PBO test on the timing rule — a coin flip. Killed before it ever touched the locked holdout.

01The pattern, and where it holds

The regime definition is a straight relative-strength trend: divide one basket’s equity by the other’s and ask whether the ratio is above or below its own trailing average. “High-Beta leading” = the ratio is above its N-day average; “Low-Vol leading” = at or below it. For the 32-year version, the two baskets are synthetic: each month, rank all point-in-time S&P 500 members by trailing 252-day beta (the standard market-model beta, correlation times the ratio of volatilities) for the high-beta decile, and by trailing 252-day volatility for the low-vol decile — bottom decile, lowest vol. Equal-weight, about 50 names each. Two mutually-exclusive long-SPY sleeves, each active only on the days its own regime is on, show what SPY actually did while each side was leading — contemporaneous, not a forecast.

The clean way to check whether that’s a real structural difference or a fluke of one lookback window is to sweep the trend window and see if the gap holds up everywhere, not just at the value I happened to pick first.

High-Beta leads on Sharpe at every lookback

SPY’s realized Sharpe while each regime is on · synthetic S&P 500 deciles, 1994–2026

High-Beta leading Low-Vol leading
Six lookbacks, 21 to 126 trading days, and High-Beta leads on every single one — a plateau, not a lucky spike at one setting. But High-Beta’s edge is in smoothness, not size: Low-Vol-leading days often earn the higher raw return (10.3% vs 8.2% annualized at N=21) — just a choppier one.

02The real signal is at the tails

Along the way, the naive high-beta decile did something extraordinary: equal-weighted, monthly-rebalanced, no cap-weighting to dilute the biggest movers, it fell 97.8% from its own peak in October 2002 — a real, mechanically legitimate result, not a data error. The stocks that measure as “highest beta” in a bear market are the ones that already crashed hardest, and an equal-weight decile keeps buying them every month on the way down. It’s the textbook betting-against-beta result: naive high-beta portfolios crash hardest exactly when beta matters most.

Growth of $1, log scale

Synthetic high-beta decile vs low-vol decile · 1994–2026

High-beta decile Low-vol decile
Both baskets end up fine — 29× and 25× over 32 years — but the paths could not look more different. The high-beta line falls off a cliff in 2001–02 and takes over a decade to fully recover; the low-vol line barely notices most drawdowns.

The crash lives almost entirely in one basket

Drawdown from prior peak · 1994–2026

High-beta decile Low-vol decile
−96% on the high-beta side (daily data touches −97.8%) against −36% on the low-vol side, and the high-beta line doesn’t make it back above −50% until 2018 — sixteen years later. That asymmetry is most of why Low-Vol so reliably leads once a real bear market gets going: it isn’t that low-vol stocks are winning, it’s that high-beta stocks are still digging out.

Split by year, the pattern is unmistakable in the crashes and absent everywhere else.

SPY’s return, split by which regime was leading that day · bear-market years
YearSPY, full yearHigh-Beta-leading daysLow-Vol-leading days
2000−9.7%−7.8%−5.3%
2001−11.8%−2.5%−9.0%
2002−21.6%−5.7%−17.1%
2008−36.8%−14.2%−27.8%
2022−18.2%−13.7%−2.2%

2000–01 read oddly because Low-Vol-leading days dominated the year (72–75% of trading days) while the market bled slowly throughout; by 2002 and 2008 the gap is unambiguous. 2022 flips the label assignment — a rate-driven growth/high-beta selloff with defensives holding up — but the same mechanism: whichever side isn’t crashing gets to “lead.”

Whichever side is crashing shows up in its own column

SPY’s return, split by regime · bear-market years, 1994–2026

2008 is the cleanest case: Low-Vol led 82% of the year, and SPY lost 27.8% on those days against 14.2% on the High-Beta-leading days. 2022 is the mirror image — a bear market that hit high-beta names hardest, so Low-Vol-leading days were the calm ones.

Outside the crashes there’s no consistent edge, and the signal visibly lags sharp turns. In the 2009 V-recovery, Low-Vol-leading days earned +26.2% against High-Beta’s +0.4% — the exact opposite of what “risk appetite” framing would predict, because the early rally days were still labeled Low-Vol from the crash momentum. COVID flips it the intuitive way: Low-Vol-leading days were the crash itself (−1.3%), High-Beta-leading days were the recovery (+6.5%). Both readings are correct for their moment; neither would have told you what to do in real time.

03So I tried to trade it

A real contemporaneous pattern isn’t the same thing as a tradable rule, so I built the one rule the finding actually implies — long SPY only while High-Beta leads, cash otherwise — and ran it through a locked-holdout validation gauntlet rather than trusting the backtest. Split the 32-year sample into a development window (1994–2016, all tuning happens here) and a locked holdout (2017–2026, touched exactly once, if at all).

On the development window, at the lookback I’d have picked going in (63 trading days), the timing rule ties buy-and-hold on risk-adjusted return instead of beating it: 5.3% a year against 9.25% for SPY, but with less than half the drawdown (−32% vs −55%). Same MAR (0.17 both), nearly the same Sharpe (0.54 vs 0.56). It cuts return and risk in roughly the same proportion — a wash dressed up as lower volatility, not an edge.

Sometimes ahead of buy & hold, sometimes behind — no pattern to it

MAR (return ÷ worst drawdown), development window 1994–2016

If the timing rule captured something real, its MAR should sit consistently above buy-and-hold’s flat 0.17 line across lookbacks. Instead it bounces both sides of it — worse at 21 and 126 days, better at 84 and 105 — the signature of noise, not structure.

04The kill

RealTest has a native overfitting test built for exactly this doubt: split the 22-year development window into 16 time blocks, form every way to pick 8 of them as “in-sample” and the other 8 as “out-of-sample” — 12,870 combinations — and ask how often the lookback that looks best in-sample also ranks well out-of-sample. A real, structural edge should win consistently across splits. Noise wins about half the time, by construction.

The result: an Overfitting Score of 51.4%. Coin-flip, and if anything a hair worse than 50% — with 12,870 splits the sampling noise on that number is tiny, so the in-sample-best lookback was actively uninformative about the out-of-sample rank, not just unhelpful. The reality haircut on top: in-sample Sharpes across splits ran 0.55–0.87, and the expected out-of-sample shrinkage was 0.21 — enough to erase most of whatever edge the in-sample number showed.

That result ends the test. The locked 2017–2026 holdout was never touched — zero peeks — because spending it on a rule that already failed the overfitting check would just contaminate data I might want for something that actually survives one day.

A real pattern in 32 years of data, and a coin flip the moment it has to pick a parameter to trade on. Both things are true at the same time.

Real crash signature, no tradable edge.

05The honest bottom line

The mechanism is real and it replicates: High-Beta-leading days are a genuinely smoother, higher-Sharpe regime for SPY, holding at every lookback from 21 to 126 days across 32 years, and every major bear market in the sample — 2000–02, 2008, 2022 — shows the same signature of Low-Vol dominating the calendar while the market bleeds. What it isn’t is a standalone timing signal: the one tradable rule it implies ties buy-and-hold rather than beating it, the parameter choice behind that rule is statistically indistinguishable from noise, and it never earned the right to spend the locked holdout that would have been the real test.

For what it’s worth, the indicator is currently reading High-Beta leading, about 10% above its own trend as of the last close in this data — a description of where the ratio sits, not a forecast, and not something to act on by itself.

Receipts

None of this makes the mechanism wrong. Betting-against-beta is one of the better-documented anomalies in finance, and low-volatility stocks really do end up “leading” when the tape is breaking, in a way that shows up cleanly across four separate bear markets and 32 years of data. It’s a legitimate thing to keep an eye on. What it isn’t — at least not in the simple long/cash form I could think to test — is something you can point a strategy at.