Data Snooping — The Trap of Overfitting
If you found a 'surefire trading rule' that fits the past 20 years of data 100%, should you celebrate or be suspicious? Data snooping is exactly the trap that makes you unsure whether that rule is real skill or coincidence.
What is data snooping?
Data snooping bias is a statistical illusion that arises when you look at the same data many times to find a 'well-fitting rule.' If you keep searching, changing all sorts of patterns and conditions until it fits the data perfectly, even a rule with no actual meaning ends up looking like a great discovery.
The core principle is simple. The more strategies you test, the higher the probability that one appears good purely by luck. This is called the 'multiple comparisons problem.' Flipping a coin once and getting heads is coincidence, but if 1,000 people each flip 10 times, someone will get 10 heads in a row. Can we call that person a 'coin-flipping genius'?
In stocks and investing, this trap is especially dangerous. Past price data is fixed, and rules (how many days for a moving average, what percent for a stop-loss, which day of the week to buy, etc.) can be combined infinitely.
Data snooping is essentially the same problem as 'overfitting' in machine learning. However, the form in which its degree can be calculated statistically is specifically called data snooping.
Overfitting: perfect in the past, collapse in the future
While backtesting (testing a strategy with past data), you become tempted to find the combination that produces the highest return by tweaking the conditions bit by bit. 'Switching from the 20-day to the 23-day line, and the stop-loss from 5% to 7%, doubled the return!' This process of whittling the rule to fit the data is exactly overfitting.
The problem is that a rule made this way memorizes even the 'noise' of the past data. Just as a student who merely memorized the answers to a practice exam collapses when facing new questions, an overfitted strategy performs far worse in the actual future market than in the backtest.
Researchers have tried to quantify this problem. Notably, the 'Deflated Sharpe Ratio' proposed by Bailey and López de Prado in 2014 shaves off inflated performance by reflecting how many attempts produced that performance. The core message is that if you don't record the number of attempts, you end up overly optimistic about the performance.
This site's calculations do not find a specific surefire rule. It only shows the actual result of buying over a long time and steadily, including the maximum drawdown and drawdown duration, and has nothing to do with predicting the future or unearthing trading rules.
A real case: the Super Bowl Indicator
The most famous case of data snooping is the U.S. 'Super Bowl Indicator.' It is a superstition that you can predict whether the stock market will rise or fall that year based on which conference (AFC/NFC) team wins the American football championship.
When journalist Leonard Koppett discovered this pattern in 1978, astonishingly it had never been wrong up to that point. The historical hit rate comes out fairly high, around 70% depending on the source. But there is no economic link whatsoever between football game results and stock prices.
Decisively, after this indicator became widely known, its accuracy fell sharply. In the 2000s (2000–2009), it was right only four out of ten times. It was a coincidence, not a real causal relationship, to begin with. Just as with 'Bangladesh butter production and the S&P 500,' any two unrelated data series, if you dig enough, can look plausibly overlapping. Correlation is not causation.
The Super Bowl Indicator's hit rate is tallied differently across sources, at about 67–75% (a difference in whether the post-application period is included). Here it is written as a range, 'around 70%.' What matters is not the exact number but that a hit rate that looks high can be coincidence.
How not to fall into the trap
The defenses used by scholars of the investment world also serve as lessons for individual investors.
First, split the data into 'in-sample' and 'out-of-sample.' Build the rule using only the earlier data, and verify it with the new data set aside for later. It's real only if it works on data that wasn't used to make the rule.
Second, honestly count the number of rules you tested. Testing a hundred and bragging only about the single best one is cheating. In academia, methods like White's Reality Check (2000) or Hansen's SPA test are used to conservatively adjust for statistical significance when comparing multiple models.
Third, you must be able to explain the reason 'why this rule should work.' A rule that fits the data with no reason is usually coincidence. That's why this site, instead of a surefire timing rule, lets you verify with actual data whether the principle itself—'good assets, long and steady'—holds.
Costs like fees and exchange rates also encourage data snooping. Leaving these costs out of a backtest makes returns look better than they really are. You have to include costs without hiding them to get a realistic result.
よくある質問
Q. Are overfitting and data snooping the same thing?
They point to nearly the same problem. Overfitting is the phenomenon where a model memorizes even the noise of past data and collapses on new data, while data snooping is the process of repeatedly digging through the same data and 'discovering' a coincidental rule as if it were real. If data snooping is the cause, overfitting is closer to its result.
Q. If the backtest return is high, isn't it a good strategy?
Not necessarily. If you keep changing conditions to fit past data, anyone can produce a dazzling backtest. What matters is whether it also works on 'data not used to make the rule,' and how many attempts produced that performance. If there were many attempts, that performance may be largely coincidence.
Q. Then isn't long-term investing also data snooping?
The key difference is 'whether there is a reason.' When you hold good assets for a long time, explainable mechanisms—corporate profit growth, dividends, compounding—are at work. In contrast, rules like day-of-the-week or sports results can't explain why they should work. If it fits the data with no reason, you should suspect data snooping.
関連ページ
📋 結果は過去のデータに基づくものです。過去のリターンは将来のリターンを保証しません。
📋 本サービスは投資アドバイスではなく、投資を理解するための教育目的で提供されています。