How Data Mining and Backtest Overfitting Mislead Investors
Try enough investment rules against past data and one of them will look brilliant. That is not evidence it works — it is what happens when you go looking, and the demonstration below takes one table to show.
Updated 2 September 2026
The chart that convinced you
Somebody shows you a chart. A simple rule, applied to the last twenty years, would have turned a modest sum into a large one. The rule is easy to describe, the data is real, and the arithmetic checks out.
Testing a rule against past prices like this is called a backtest, and there is nothing wrong with doing one. The problem is what happens before you are shown the result.
If somebody tries fifty rules and shows you the best one, the winner is not the best rule. It is the luckiest. And nothing about the chart tells you how many others were tried.
Watching it happen
This is easier to believe when you see it, so here is the experiment run on real Indian market data.
Take a family of very ordinary rules. Each one says: stay invested while the market's price is above its own average over recent months, and move to cash while it is below. The only difference between them is how many months go into that average — 5 versions, nothing clever.
Now split the record in half at 7 January 2013. Pick the best-performing rule using only the first half. Then take that rule, change nothing about it, and run it on the second half — years it has never been tested against.
| Rule | Years it was chosen on | Years it had never seen |
|---|---|---|
| 50-day average | +12.52% | +7.42% |
| 100-day average — the winner | +14.98% | +7.92% |
| 150-day average | +11.50% | +7.02% |
| 200-day average | +9.58% | +7.02% |
| 250-day average | +10.86% | +7.51% |
| Doing nothing and staying invested | +14.39% | +12.10% |
In the first half the winner was the 100-day version, returning +14.98% a year. Simply staying invested returned +14.39%. On that evidence the rule looks like it has found something real.
In the second half, that same rule returned +7.92% a year. Staying invested returned +12.10%.
Nothing was adjusted between those two columns. No setting was re-tuned, no rule rewritten. The only difference is that the first column had already seen the data and the second had not.
Where the illusion comes from
Trying lots of things. With 5 rules the winner is chosen partly for being good and partly for being lucky. Try a hundred and the winner is almost entirely luck. There is no warning sign — the successful chart looks identical either way.
Deciding what to test after glancing at what did well. Choosing to test Indian technology shares because you already know they did well is the same mistake wearing a different hat.
Using information that was not available at the time. Economic figures are revised. Company results are restated. The list of companies in an index today is not the list from ten years ago. All of these leak later knowledge into an earlier test.
Leaving out what disappeared. A test run only on companies that still exist has quietly deleted every failure.
Choosing convenient start and end dates. A test that begins just after a crash and ends just before the next one is a statement about two dates, not about a strategy.
Adjusting the rule after seeing it fail. Each tweak fits the past a little more snugly and the future a little less.
What a careful person does instead
- Write down the rule and the reason for it before testing. If the explanation only appears after the result does, it is a story about the data rather than a reason to expect anything.
- Set aside part of the data and genuinely do not look at it. Once.
- Charge realistic trading costs. The rules in the table above trade frequently, and their figures exclude costs entirely — which flatters them and does nothing for the "doing nothing" row.
- Check the rules either side of the winner. If the 100-day version works and the ones just above and below it do not, you have found a fluke rather than a mechanism.
- Compare against the simple option. Here that is staying invested, and it is a far harder benchmark than most strategies admit.
- Ask what else was tried. One result reported out of many attempts is not a finding.
Even a holdout is not proof
The obvious defence is to keep some data back and test on it at the end. That helps. It is not a guarantee.
Check the held-back data, adjust the rule, check again — and it has stopped being held back. This happens gradually, one reasonable-looking step at a time, without any single decision feeling dishonest. After enough rounds it is simply more of the data you already fitted to.
A genuine pattern can also stop working. Markets change: more people notice the same thing, costs fall, rules change. A relationship can be real, properly tested, and still expire.
Three questions to ask of any backtest
- How many versions were tried before this one was chosen?
- What did the ones that failed look like?
- What would have had to happen for this to be reported as a failure?
If those answers are not available, the chart is not evidence. The table above is what it looks like when the second column finally arrives.
Educational content only. This is not personalised financial, investment or tax advice. Figures quoted are historical or illustrative and are not forecasts. Consult a qualified professional before acting on anything you read here.