Kairon AI
Backtesting & AI

Backtesting pitfalls: why most track records mislead you

Overfitting, look-ahead bias, survivorship bias, cherry-picking and tiny samples, explained with examples, plus how to read an out-of-sample, walk-forward result.

Updated 2026-09-25 · 3 min read · Kairon AI Research

If you have seen an ad for a trading system, you have seen an equity curve that goes up smoothly from left to right. Almost all of them are real, in the sense that someone ran the numbers. Almost all of them are also misleading. This guide explains why, and what an honest result looks like.

1. Overfitting: tuning to the noise

Give a computer enough knobs and enough history and it will find settings that would have made money. A moving-average crossover with a 17-day and 43-day window, applied only on Tuesdays, might look excellent on 2019 to 2024. It is not a discovery. It is a description of noise.

The more parameters a strategy has and the more combinations were tried, the more likely the result is fit to the past rather than predictive of the future.

How to spot it: very specific parameters, dramatic results, and no separate test on unseen data.

2. In-sample vs. out-of-sample

  • In-sample: the data used to build and tune the rules.
  • Out-of-sample: data the rules never saw during tuning.

In-sample performance is optimistic by construction, because the rules were chosen to look good on exactly that data. Only out-of-sample performance is an estimate of what might happen next.

3. Walk-forward testing

The standard way to get out-of-sample results from history:

  1. Tune the rules on a first period (for example the first 12 months).
  2. Test them, frozen, on the next period (the following 3 months).
  3. Move both windows forward and repeat.
  4. Report only the combined results of the test periods.

Each test period is out-of-sample for the rules used in it. This is what Kairon uses for the figures it labels out-of-sample. It is also why our out-of-sample numbers are lower than our in-sample numbers, which is exactly what you should expect from an honest test.

4. Look-ahead bias

The simulation uses information that was not available when the decision was made. Common examples:

  • Trading at the open with a signal that uses that day's close.
  • Using financial data as later revised instead of as first reported.
  • Using an index list from today to test a strategy in 2015.

It is subtle, usually unintentional, and it makes results look much better. Good backtesting code has explicit guards against it.

5. Survivorship bias

Testing only on companies that exist today silently removes every bankruptcy, delisting and takeover at a low price. A strategy that "buys beaten-down stocks" looks brilliant if the ones that went to zero are not in the data.

6. Cherry-picking and small samples

  • Cherry-picking: showing the best period, the best asset or the best variant out of many. Ask how many variants were tested.
  • Small samples: 25 trades can produce a 70 % win rate by luck alone. As a rough rule, be sceptical of any claim built on fewer than 100 to 200 independent trades, and of any single market regime.

7. Ignoring costs and execution

Fees, spreads, slippage and gaps are small per trade and large in total. A strategy that trades often can turn from profitable to losing once realistic costs are included.

How to read a track record honestly

QuestionGood signWarning sign
Is every trade published?Full list, including losersOnly highlights
In-sample or out-of-sample?Clearly separatedOne blended number
Sample size and period?Stated, 100+ tradesMissing or tiny
Win rate or expectancy?Both, plus profit factorWin rate only
Costs included?Yes, statedNot mentioned
Method explained?Entry, exit, stops, holding period"Proprietary"

What we do with our own numbers

Kairon publishes every graded call, logs the method (entry at the next available price, target, stop, breakeven rule, maximum holding period, a flat zone of ±1.5 %), separates in-sample from out-of-sample, and excludes low-confidence calls from grading because they are not tradeable signals. When our numbers get worse, they get worse in public. You can check all of it in the track record.

That will not make every number look good. It will make every number mean something.

Kairon AI

Get a second opinion on your next stock

Five AI agents look at technicals, fundamentals, news and sentiment, then a bull and a bear argue it out. A free account includes one compact AI analysis every month.

Start a free analysis No credit card. Research tool, not financial advice.

This guide is educational and not investment advice. Past and simulated results do not guarantee future results.

Kairon AI

Get a second opinion on your next stock

Five AI agents look at technicals, fundamentals, news and sentiment, then a bull and a bear argue it out. A free account includes one compact AI analysis every month.

Start a free analysis No credit card. Research tool, not financial advice.

Keep reading