P-Hacking and Multiple Testing
How trustworthy is the phrase 'statistically significant'? If you repeated the test hundreds of times until a significant result appeared, that significance may be chance rather than truth.
What P-Hacking Is
In statistics, the p-value roughly shows 'the probability that this result appeared purely by chance.' Conventionally, a p below 0.05 is called 'statistically significant' — meaning it's too rare to be chance.
P-hacking is the practice of forcibly manufacturing a 'significant' result by repeatedly testing — changing variables, periods, and samples this way and that — until you cross that 0.05 threshold. It's also called 'data dredging.'
If it doesn't work at once, change the conditions a little and try again, change them again and try again — test about twenty times and one of them will come out with p < 0.05 purely by luck. Pick just that one and announce 'a significant finding!' and you've actually packaged chance as truth.
The American Statistical Association (ASA) warned in its 2016 statement that if you do data dredging and multiple testing, 'the reported p-value becomes essentially uninterpretable.'
The Trap of Multiple Testing
The key is 'the number of tests.' A p = 0.05 threshold means 'the probability of coming out significant by chance is 1 in 20.' So if you test 20 different hypotheses, on average 1 of them will come out 'significant' by chance even if it means nothing. This is the multiple comparisons problem.
Test 100 or 1,000 and the chance 'significant findings' pile up accordingly. The problem is that when publishing, many don't reveal 'how many they tested to pick this one.' The reader loses any basis to judge whether the result is real or luck.
Statistics has methods to correct for this. The representative Bonferroni correction makes the threshold stricter in proportion to the number of tests. For example, if you test 20, lower the threshold from 0.05 to 0.05 ÷ 20 = 0.0025 to filter out chance significance.
When you see a claim of 'significance,' always ask: 'How many did they test to find this?' Significance without disclosing the number of tests should be treated with caution.
What Happens with Investment Factors
This problem is especially severe in investment research. Academia and industry have announced countless 'factors' claiming that 'follow this metric and you'll beat the market.'
The financial economist Campbell Harvey's research team, in a 2016 paper, cataloged as many as 316 factors published by academic journals. They pointed out that if this many factors were dug out of the data, a large share of them must be false discoveries that 'looked to work' by chance rather than real effects.
So the team proposed that for a new factor to be accepted, it must clear not the usual threshold (t-statistic 2.0) but a far stricter t-statistic of 3.0 or higher. They raised the bar to account for the sheer number of tests. In other words, a large share of factors that 'worked in past data' may in fact be products of p-hacking, dug out by chance after hundreds of tests.
Harvey, Liu & Zhu (2016), '... and the Cross-Section of Expected Returns,' showed statistically that a large share of existing factor studies that didn't correct for multiple testing are likely false discoveries.
Why Simple Rules Are Strong
The opposite of p-hacking is 'not selecting the best by changing conditions this way and that.' The fewer tests you run and the simpler the rule, the less room there is to mistake chance for truth.
This is also why The Return of Almost Everything doesn't boast complex 'secret factors.' We aren't trying to show that 'this combination of metrics beat the market in the past'; we show, exactly as it was — including maximum drawdown and loss periods — how much you'd have if you had bought a good asset over a long, steady period. There's no process here of 'changing conditions until it becomes significant.'
When you look at numbers, ask: 'How many times did they test to find this result? Did they disclose the number of tests?' This one question helps you avoid being fooled by a result that merely sparkled by chance.
The simpler the rule, the less room to secretly fit it. The simple principle of 'long-term and steady' surviving longer than dazzling factors is partly for this reason.
常见问题
Q. If the p-value is below 0.05, can I trust it unconditionally?
No. p < 0.05 is only a weak signal that 'it's rare to be chance,' and if the test was repeated many times, its meaning weakens greatly. The American Statistical Association also warned in its 2016 statement that a single p-value near 0.05 is only weak evidence, and that with multiple testing the p-value becomes essentially uninterpretable.
Q. How do I correct for multiple testing?
The simplest method is the Bonferroni correction. Divide the significance level by the number of tests to make the threshold stricter. For example, if you test 20 hypotheses, lower the threshold from 0.05 to 0.0025. In investment factor research, some, like Campbell Harvey's team, raise the required t-statistic itself from 2.0 to 3.0 or higher.
Q. So are all claims that 'this metric beats the market' false?
Not all are false, but you must view them very cautiously. Digging through countless metrics in data will inevitably turn up some that happened to work in the past by chance. The more a factor fails to disclose 'how many were tested' and 'whether it holds even in periods never looked at,' the safer it is to keep in mind that it may be a product of p-hacking.
📋 结果基于历史数据计算,过去的收益不代表未来的收益。
📋 本服务旨在帮助理解投资、供教育之用,并非投资建议。