ACADEMY ·  Tools & Strategies ·  Indicators & Tools
Indicators & Tools  ·  Lesson 16 of 20

Backtesting an Indicator Rule Honestly

A hand-testing protocol: sample size, market states, and the pitfalls (hindsight, curve-fit) that fake good results.

5 MIN READ · THE DESK ACADEMY

A trader posts a chart showing a MACD crossover rule that won 13 of 15 trades over two weeks and calls it a backtested edge. It is not a backtest. It is fifteen coin flips dressed up with a formula, and the sample is too small to tell a real edge from luck at any confidence worth acting on. Honest testing needs a real sample, real variety in market conditions, and a process that refuses to peek at the answer before writing down the question.

The sample size that actually means something

A rule needs on the order of 100 or more signals before its win rate and average result mean much of anything, because smaller samples are dominated by variance rather than edge. A rule tested on 15 or 20 trades can look brilliant or terrible almost at random, and both readings will feel completely convincing while they last. Building toward 100 clean signals usually means testing across several months of a chosen instrument's history, not a single lucky fortnight.

Testing across market states, not just time

More history is not the same as better history if it all comes from one kind of market. A rule tested only across a trending quarter will look like a moneymaker and then fall apart the first time it meets three weeks of chop, and the reverse is just as common for range based rules meeting a trend. Pick a testing window deliberately covering both a clearly trending stretch and a clearly ranging stretch on the instrument you actually trade, and report the results separately for each, not blended into one flattering average. A rule that performs acceptably in both states is worth far more than one that shines in a single regime and says nothing about the other.

Take a 20 EMA pullback rule tested on Nasdaq across 120 signals over three months. Split cleanly, 70 of those signals fell inside a trending month and 50 fell inside a choppier one. Reporting a blended 58 percent win rate across all 120 hides the more useful fact: the rule won 68 percent of the time during the trending month and only 44 percent during the choppy one. That split tells a trader exactly when to lean on the rule and when to sit on hands, which the single blended number never could.

The hand testing protocol

Scroll the chart forward one bar at a time rather than looking at the whole history at once. At each bar, decide only from what would have been visible at that moment whether the rule's conditions are met. Record the entry, the stop and the target before scrolling forward to see what happened, exactly as if the trade were live. This single habit, writing the plan down before seeing the outcome, is what separates genuine testing from scrolling backward through a chart where hindsight quietly edits every decision.

The two pitfalls that fake good results

Hindsight bias creeps in the moment you test while looking at the full chart at once, because your eye already knows which signals worked and unconsciously favors marking those as valid entries. Curve fitting creeps in when a rule that performs poorly gets its settings nudged, an RSI period changed from 14 to 9, a moving average from 20 to 17, until the same historical data finally produces a good result. A rule tuned that precisely to one stretch of the past rarely survives contact with the future, because it was never finding an edge, it was memorizing noise.

Knowledge pays better with capital behind it.

Practice this on a free $10K account, or trade a Daily Funded Session where a disciplined, profitable day pays out the same day.

Start a Funded Session