A Dow Theory rabbit hole · Sep 2026
I’ve heard about Dow Theory for years — the Industrials and the Transports have to confirm each other before a trend is real — and never once seen anyone simply backtest it. So I turned the rules into code and ran them across 126 years of Dow data. Does confirmation add anything? Does it beat luck? And does it still work today?
Credit where it’s due: this is the grandparent of every trend-following system, and the core idea is genuinely clever. Charles Dow, and later William Hamilton and Robert Rhea, argued that the economy has two halves that have to agree. Factories make things; railroads move them. If the Industrials rally to a new high but the Transports don’t follow, the goods aren’t actually shipping — and the rally is suspect. It’s a real economic intuition, a century before anyone said “cross-asset confirmation.”
The reason nobody backtests it is that Dow Theory is written as judgment, not rules: “secondary reactions,” “lines,” “confirmation.” Any test has to choose a mechanical reading. So I did two things to keep myself honest: I picked a plain, textbook reading before looking at results, and then I ran ten variations of it and report all of them, not the winner.
Return vs buy & hold
Average across 10 variants, 1900–2026: 8.8% vs 9.8% a year. Only 3 of the 10 beat buy & hold. The base rule ties it.
Worst drawdown
Base rule, vs −87% for buy & hold. Most of that is 1929–32. The 10 variants average −48%.
Beat buy & hold since 1953
On return. Drawdown is still ~13 points shallower (−39% vs −52%), so it buys safety, not extra return.
This is the most common modern reading of Hamilton’s version. For each index, I track the swing highs and lows using closing prices only. A secondary reaction is a pullback (or bounce) of at least 5% from the running extreme. Then:
Nothing in the rule can see the future: each day only uses swings that were already confirmed. Here is what it did to the Dow over 126 years. Shaded bands are the stretches when it said “cash.”
Dow Jones Industrial Average, log scale · shaded = Dow Theory in cash
Over the full sample the base rule earns 9.8% a year, the same as buy and hold — but with a worst drawdown of −54% instead of −87%, and only 72% of the time in the market. That’s a genuinely better ride. Before getting excited, though, that base rule sits at the high end of my ten variants; the honest number is the average (8.8%), a point below buy and hold.
| Strategy | CAGR | Max DD | MAR | Sharpe* | In mkt | Trades |
|---|---|---|---|---|---|---|
| Buy & hold the Dow | 9.8% | −87.0% | 0.11 | 0.61 | 100% | — |
| Dow Theory · base rule (5%, both indexes) | 9.8% | −54.0% | 0.18 | 0.80 | 72% | 76 |
| Dow Theory · average of all 10 variants | 8.8% | −48.2% | 0.19 | 0.72 | — | — |
| Industrials only, no confirmation (avg of 10) | 8.3% | −61.8% | 0.14 | 0.68 | — | — |
| Transports only, no confirmation (avg of 10) | 9.0% | −44.9% | 0.21 | 0.76 | — | — |
| 200-day moving average (Dow) | 9.1% | −54.1% | 0.17 | 0.83 | 65% | 424 |
| Same 72% exposure, random timing (blend) | 8.5% | −75.1% | 0.11 | 0.69 | 72% | — |
*Raw return divided by volatility, no cash rate subtracted. Measured against cash, buy & hold is 0.41, Dow Theory 0.51 and the 200-day average 0.50. MAR = CAGR divided by worst drawdown.
Two things stand out. First, the 200-day moving average does about the same job — slightly lower return, the same drawdown, about the same Sharpe — without needing a second index. Second, the “blend” row is a useful reality check: simply holding 72% stocks and 28% cash, with no timing at all, earns 8.5% but still takes a −75% drawdown. Being in cash at the right times is what the Dow rule adds.
Total return, log scale · next-close fills, 0.10% per switch
Drawdown from prior peak · 1900–2026
Split the 126 years into eras and the picture changes. The rule was excellent in the early decades — +3 points a year over buy and hold from 1900 to 1929 — and roughly break-even after 1950 on return, with shallower drawdowns until 2000. Since 2000, it has lagged buy and hold by two points a year, and its drawdown isn’t any better.
| Period | B&H CAGR | Dow CAGR | B&H max DD | Dow max DD | In mkt |
|---|---|---|---|---|---|
| 1900-1929 | 10.2% | 13.1% | −47.5% | −31.3% | 66% |
| 1930-1949 | 4.6% | 3.9% | −83.5% | −48.5% | 58% |
| 1950-1979 | 9.2% | 9.3% | −41.1% | −35.3% | 76% |
| 1980-1999 | 17.9% | 17.2% | −35.9% | −21.4% | 83% |
| 2000-2026 | 8.2% | 6.2% | −51.8% | −50.0% | 77% |
Annualized total return by era
The episode table shows why. In the classic crashes it did what the theory promises. In the last two big drawdowns that were choppy rather than one-way, it did the opposite.
| Episode | Buy & hold | Dow Theory |
|---|---|---|
| 1929–32 crash (Sep ’29–Jul ’32) | −82.8% | −43.8% |
| 1937–38 recession | −33.0% | −6.2% |
| 1987 crash (Aug–Dec) | −23.6% | −11.3% |
| 2000–02 dot-com bear | −23.4% | −47.9% |
| 2007–09 financial crisis | −42.8% | −29.1% |
| 2020 COVID crash (Feb–Mar) | −22.0% | −4.1% |
| 2022 rate shock (Jan–Oct) | −8.4% | −19.0% |
The 2000–02 bear is the instructive one. The Dow fell nearly 40% peak to trough — but in a series of violent rallies. The rule sold in February 2000, October 2000, March 2001 and September 2001, and each time it re-bought after a bounce, 4% to 16% above where it had sold. Four sells, four higher re-buys. 2022 rhymed: it sold in September near the low and re-bought 15% higher in February. In the 1987 crash, by contrast, it flashed a sell on October 15 and was in cash for Black Monday.
This is the question the whole theory rests on. If you need both indexes to agree, you should do better than watching either one alone. So I ran the same ten parameter variants three ways: Industrials only, Transports only, and both required to confirm.
Average annualized return across 10 parameter variants · 1900–2026
So the “two indexes must agree” rule adds something over the Industrials, but it isn’t clear it adds anything over just watching the Transports. That’s a plausible story — the Transports are the more cyclical, more volatile index and tend to roll over first — but I’d treat it as a hypothesis, not a finding. It could equally be that a faster, noisier index just makes a better trend filter.
The parameter grid shows how much the result depends on the choices. Every variant is below, with the base rule highlighted:
| Reaction size | Min duration | CAGR | Max DD | MAR | Buys |
|---|---|---|---|---|---|
| 3% | none | 9.2% | −43.4% | 0.21 | 152 |
| 3% | 10 days | 10.0% | −45.2% | 0.22 | 91 |
| 5% | none | 9.8% | −54.0% | 0.18 | 76 |
| 5% | 10 days | 10.5% | −43.5% | 0.24 | 55 |
| 7.5% | none | 8.5% | −59.2% | 0.14 | 43 |
| 7.5% | 10 days | 8.9% | −39.3% | 0.23 | 36 |
| 10% | none | 7.2% | −49.5% | 0.15 | 30 |
| 10% | 10 days | 7.8% | −47.5% | 0.16 | 26 |
| 15% | none | 7.8% | −53.0% | 0.15 | 12 |
| 15% | 10 days | 8.4% | −47.5% | 0.18 | 11 |
Nothing here is a cliff-edge, but the trend is clear: CAGR is about 9–10.5% for 3–5% reactions and drops to 7–8.5% for 10–15%, because bigger thresholds react later. Adding the 10-day minimum duration helps in every row, by 0.4 to 0.8 points. The base rule is not the best cell (that is 5% with the duration filter, 10.5%) but is third best of ten.
A fair worry: with only 76 trades, could any in-and-out pattern with the same amount of cash time look this good? I tested it directly. I took the rule’s exact in/out pattern — same number of trades, same lengths of time in cash — and slid it to a random position along the 126 years, 3,000 times. If the timing were meaningless, the real result would land in the middle.
Return per unit of worst drawdown (MAR), 3,000 random shifts of the same trade pattern
Two caveats keep that from being a victory lap. First, the random shifts still carry the 1930s inside them, so part of what “real” means here is that the rule reliably caught the big early crashes. Second, individual sell signals are right less often than a coin flip: of 75, only 31 (41%) were followed by lower prices, and 44 by higher. The value comes from size, not frequency. The six best calls account for about half of all the decline the rule avoided.
The most-cited modern mechanical version is Jack Schannep’s. His published record claims 13.7% a year for 1953–2025 against 10.9% for buy and hold. Its rules differ from mine: it adds the S&P 500 as a third index, sets a 3% minimum reaction with duration tests, and adds exceptions such as a shorter duration after a capitulation. I can’t audit that record. What I can say is that the plain two-index version doesn’t get there. Its closest cousin in my test (3% reaction, 8-day minimum) earned 8.4% a year since 1953 against 10.7% for buy and hold, with a max drawdown of −29% (vs −52%). Better risk, lower return.
I also tried shorting the sell signals instead of moving to cash. It’s a disaster: 6.9% a year with a −76% drawdown. The bear phases are too often followed by rallies.
Dow Theory is real, but it’s a smaller thing than its reputation. It isn’t a return booster: across ten variants it trails buy and hold by about a point a year over 126 years, and by nearly two since 1953. What it is is a slow, crash-avoiding trend filter — it beat random timing at the 99th percentile, kept you out of most of 1929–32, 1937, 1987 and 2020, and cut the worst drawdown by a third. It works best when markets fall in one long slide, and worst in choppy bears, where it sells low and re-buys high (2000–02, 2022). The 200-day average does nearly the same job with one index instead of two.
For what it’s worth, the rule is currently long: its last signal was a buy on May 2, 2025, with the Dow near 41,300, and no bear confirmation has followed. That’s a description of the model’s state, not a forecast.
$DJI) and Transportation Average ($DJT) daily closes from Norgate Data, 1896–2026; test starts 1900 after a warm-up. Signals use price closes only.$DJITR) from late 1987. Before that, price return plus an S&P 500 dividend yield (Shiller data). That proxy runs about 0.5 points a year below the real index over 1988–2023, so early-era returns are slightly conservative for every strategy alike.M13002US35620M156NNBR) before 1934, 3-month T-bill (TB3MS) after.None of this makes the old idea wrong to respect. Two indexes that have to agree is a good way to make a trend prove itself, and the crash protection in 1930 and 1987 was real. Chasing it one rabbit hole deeper just changes what to expect from it: not a way to beat the market, but a way to sit out some of its worst stretches — at the price of some whipsaws and a good deal of patience.