Multiple Testing Problem
The multiple testing problem (also called the multiple comparisons problem) is the statistical fact that running many tests on the same data inflates the chance of finding a false positive. Each individual test carries a small risk of flagging a chance result as meaningful, and those small risks add up. Test enough ideas, and some will appear significant even when nothing real is there.
This problem is the engine behind several related research hazards in quantitative finance. It explains why data snooping manufactures spurious patterns, why p-hacking works, and why so many proposed return drivers in the factor zoo fail to replicate. Understanding the mathematics of multiple testing clarifies why a single impressive result, drawn from a wide search, deserves heavy skepticism.
Definition
A standard significance test is designed to keep the false-positive rate of one test low, often at 5%. That means that even when no real effect exists, the test will wrongly flag significance about 1 time in 20. Running the test once is fine. Running it many times on the same data multiplies the opportunities for a chance result to slip through, so the probability that at least one test produces a false positive climbs quickly with the number of tests.
Key Principle
With a 5% threshold and 20 independent tests, the chance that at least one false positive appears is not 5% but roughly 64%. At 100 tests it exceeds 99%. The threshold that protects a single test offers almost no protection across a large search, which is why the number of tests must be accounted for explicitly.
Why It Matters in Finance
Quantitative finance is unusually vulnerable to the multiple testing problem because researchers commonly screen hundreds or thousands of candidate signals against the same historical data. The t-statistic (a standardized measure of how far a result sits from zero) of a single backtest can look convincing in isolation, yet that same threshold is far too lenient once you account for the many strategies that were tried and discarded.
Harvey, Liu, and Zhu (2016) argued that, given decades of factor research mining the same data, the conventional two-standard-deviation cutoff is too low. They proposed substantially higher hurdles to declare a factor genuine. The implication is direct: a result that would pass an ordinary test may fail once the breadth of the search is properly counted.
Common Corrections
Statisticians have developed several adjustments that raise the bar for significance in proportion to the number of tests. They differ in how strict they are and in exactly what they try to control.
| Method | What It Controls | Character |
|---|---|---|
| Bonferroni correction | The chance of any false positive across all tests | Simple and strict; can be overly conservative with many tests |
| Holm step-down | The same family-wide error rate, less conservatively | More powerful than Bonferroni while still controlling false positives |
| Benjamini-Hochberg (FDR) | The expected proportion of false positives among discoveries | Less strict; better suited to screening many candidates |
The simplest of these, the Bonferroni correction, divides the significance threshold by the number of tests. If 20 strategies are tested, each must clear a 0.25% threshold rather than 5% to count as significant. This controls false positives effectively but grows very strict as the number of tests rises, which is why methods that control the false discovery rate are often preferred when screening large numbers of candidates.
Known Limitations
Limitations to Keep in Mind
- The number of tests is often unknown. Corrections require knowing how many hypotheses were examined, but researchers lose count, and the field-wide total across everyone studying the same data is impossible to measure. The adjustment is therefore an estimate, not an exact fix.
- Strict corrections cause false negatives. Raising the bar to suppress false positives also rejects some genuine effects as noise. There is an unavoidable trade-off between missing real discoveries and accepting false ones.
- Tests are usually correlated. Many corrections assume the tests are independent, but financial strategies often overlap heavily. Correlated tests make the true error rate harder to compute, so the standard formulas are approximations.
- Corrections cannot fix selective reporting. If only the winning tests are disclosed and the failures hidden, no correction can recover the truth, which links the problem to publication bias. Honest accounting of all tests run is a prerequisite.
Further Reading
- Harvey, C.R., Liu, Y. and Zhu, H. (2016). "...and the Cross-Section of Expected Returns." The Review of Financial Studies, 29(1), 5–68.
- Benjamini, Y. and Hochberg, Y. (1995). "Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing." Journal of the Royal Statistical Society: Series B, 57(1), 289–300.
- Harvey, C.R. and Liu, Y. (2020). "False (and Missed) Discoveries in Financial Economics." The Journal of Finance, 75(5), 2503–2553.
Related Terms
Foxholm Financial is a fee-only registered investment adviser serving Georgia. We bring quantitative rigor to every client engagement. Explore our services or get in touch to discuss how we can help. To see how this kind of analysis informs real client work, explore a Strategic Portfolio Review.
Are you an institution or FinTech firm? Learn about our Quantitative Consulting Services.
Foxholm Financial trains the next generation of quantitative analysts. Students and early-career researchers can explore our quantitative investment fellowships.
This content is for educational and informational purposes only and does not constitute an offer to sell or a solicitation of an offer to buy any securities. Nothing herein constitutes investment advice or recommendations tailored to your individual situation. All investments involve risk, including the potential loss of principal. Past performance is no guarantee of future results. Information presented is believed to be factual and up-to-date, but Foxholm Financial does not guarantee its accuracy and it should not be regarded as a complete analysis of the subjects discussed. Before making investment decisions, consult with a qualified financial advisor who can evaluate your specific circumstances.