What is backtesting?
Backtesting is the process of applying a defined trading strategy to historical observations to simulate how its rules might have behaved. The researcher specifies the conditions, timing, position logic, and assumptions, then evaluates the decisions those rules would have produced in the selected data.
A backtest is an experiment with a model of past conditions, not a record of actual trades and not a forecast. Its meaning depends on what data and execution assumptions were used. For the full research sequence around a systematic idea, see the Algorithmic Trading Workflow and the beginner’s guide.
What a backtest is designed to test
A strategy-level backtest can help answer whether rules are internally coherent, how often specified conditions occurred, and how the simulated process behaved under a declared set of assumptions. It can expose unintended interactions, such as an entry condition that repeatedly triggers while a position is already open, or an exit rule that depends on information not available at the decision time.
It does not prove that the underlying hypothesis is true, that the process can be executed at modeled prices, or that a historical relationship will persist. A backtest also does not test every part of a running trading application. Model-validation methodology belongs to AI Trading Model Testing, while end-to-end software, integration, and recovery checks belong to Testing Trading Bots.
Historical data requirements
The data should match the question and the intended frequency of decisions. Daily observations may be suitable for a slower research question but insufficient for a process whose assumptions depend on intraday order or price changes. Review source coverage, timestamps, units, instrument identifiers, missing records, and any transformations applied.
Consider whether the historical universe reflects instruments that were actually eligible at each point in time, rather than only those available today. Where relevant, corporate actions and changes to symbol identifiers should be handled consistently. Data errors or incomplete coverage can affect signals and simulated outcomes before execution assumptions are considered. The deeper guide to historical data integrity for systematic trading covers point-in-time and data-consistency issues.
Define the strategy precisely
The simulator can only evaluate the rules it is given. Specify the entry and exit conditions, when each is checked, what happens when several conditions are true, how positions are sized, and whether additional positions can be opened. State the eligible instruments, timeframe, and any restrictions that shape the process.
Make the decision timeline explicit. If a condition uses a closing value, establish whether a simulated order is placed at a later observation rather than at a price that was known only after the decision. If rules are expressed in vague terms, different implementations may produce different tests. Precise definitions help readers separate a strategy idea from the assumptions introduced by a particular backtest.
Transaction costs
A simulated result before costs may not represent the economics of repeated transactions. Depending on the market and instrument, relevant costs can include commissions, exchange or venue fees, taxes, financing, borrow, or other charges. Which apply varies, so the test should state what was included and what was not.
Costs can accumulate differently depending on turnover: a process that changes positions often may be more sensitive to per-trade assumptions than one that trades infrequently. Do not add unsupported estimates merely to make a backtest appear realistic; document their source or treat them as uncertain assumptions and examine how conclusions change.
Spread and slippage
The quoted or recorded price is not necessarily the price at which a hypothetical order could be filled. The spread is the difference between available buy and sell prices at a moment; slippage describes the difference between an assumed or requested price and the realized execution price. A backtest may need a simple approximation, detailed historical quotes, or a range of scenarios depending on its purpose and available data.
If the test assumes every order transacts at a midpoint, opening value, or observed bar price without justification, it may understate execution friction. Spread and slippage vary with instrument, size, liquidity, timing, and market conditions. A single fixed assumption should not be treated as universal.
Timing assumptions
Timing determines what information the strategy could use and when a simulated action could occur. Establish whether inputs are known at the start or end of a period, how events with different timestamps are aligned, and how delayed or missing observations are treated. The simulation should not allow a decision to use a value before it would have been available.
Price bars summarize activity within intervals and may not reveal the order in which intraperiod prices occurred. If a strategy’s stop and target could both be reached within one bar, a bar-based test may not know which came first without finer-grained information or an explicit convention. State the limitation rather than choosing whichever order makes the result look more favorable.
Position sizing and risk assumptions
Position sizing affects the path of simulated account value and exposure. Document whether size is fixed, based on account value, constrained by available capital, or adjusted under a defined rule. Also specify what happens when the intended size exceeds a stated limit or when positions overlap.
Risk assumptions may include maximum exposure, leverage, instrument restrictions, or how simultaneous positions are handled. These are part of the test definition, not optional details that can be inferred from a headline return. Backtest assumptions are not recommendations for appropriate position sizes or risk limits.
Avoid look-ahead bias
Look-ahead bias occurs when the backtest uses information that would not have been available at the simulated decision time. This can happen directly, such as using a future price in a rule, or indirectly when a timestamp, data publication delay, or calculation convention is misunderstood.
For each input, ask when the value was observable and when the strategy could reasonably have acted on it. Rolling calculations must use only the information available up to the decision point. News, economic data, and revised datasets need availability timestamps rather than only the period they describe. Review the full path from source to feature to decision; a historical label does not by itself prove that a value was knowable then.
Avoid survivorship bias
Survivorship bias can occur when a historical test includes only instruments that remain available or prominent at the end of the period. Failed, delisted, merged, or otherwise removed instruments may be absent, making the tested universe different from the one a researcher could have selected historically.
The effect depends on the research question, data, and universe construction. Document how instruments enter and leave the sample and whether historical membership information is available. A test on a current list should not be described as though it necessarily represents the full historical opportunity set.
Overfitting and excessive parameter tuning
When many variations are tried, some may appear unusually favorable by chance. Adjusting thresholds, periods, filters, instruments, and rules repeatedly in response to the same historical results can overfit the research process even if the final strategy is simple.
Keep a record of experiments, state which choices were made before and after examining results, and avoid presenting a repeatedly tuned period as independent confirmation. A simple baseline can help establish whether added complexity changes the behavior meaningfully. This article addresses strategy-level historical evaluation; detailed AI/ML model splitting, leakage, and selection issues are discussed in the AI model testing guide.
In-sample and out-of-sample testing
In-sample observations are used while exploring or developing a strategy. Out-of-sample observations are held apart from that development and used as a later check. Evaluating a fixed rule on data not used to design it can provide more informative evidence than repeatedly testing and tuning on one period.
The separation only helps if it is respected. If the researcher changes a rule after viewing the supposed holdout result, that result has influenced development. A fresh evaluation period may then be needed, and even a genuinely separate historical test cannot guarantee future performance. This introductory distinction does not replace model-specific validation methods.
Interpreting results
Interpret multiple aspects of behavior rather than focusing on one return number. Return describes a change under the simulation’s assumptions, while drawdown describes a decline from a prior simulated peak. Volatility reflects variability in observed outcomes under a chosen measurement. Trade count and turnover help show how frequently the process acts or changes exposure; costs indicate how assumptions about transactions affect results.
Robustness is not one score. It concerns whether conclusions depend on a narrow time period, instrument set, parameter choice, or optimistic execution assumption. Examine concentration, periods of loss, and sensitivity to reasonable alternatives. Metrics need clear definitions and context; none proves that a strategy is sound, suitable, or likely to perform in the future.
Why backtests can mislead
A backtest can mislead if data is incomplete, timing is wrong, costs are omitted, the tested universe excludes historical failures, rules were repeatedly selected on the same sample, or execution assumptions are implausible. It can also mislead when the report hides negative periods, uncertainty, or the number of alternatives explored.
The appropriate response is not to make every simulation maximally complicated. It is to match the test detail to the question, disclose simplifications, and avoid claiming more than the evidence shows. If a tested strategy moves into software, system behavior needs its own testing; simulated strategy performance cannot verify order handling or recovery.
A hypothetical example
Suppose a researcher defines a rule that identifies a condition using the closing observation of each period and then records the following period’s outcome. A careful test would first confirm that the close is not used to claim an execution at a price that was only known after that close. It would document which instruments were included at each historical date, how missing observations were handled, and whether relevant fees and a spread assumption were modeled.
The researcher might compare the specified rule with a simple baseline and inspect trade count, turnover, drawdowns, variability, and performance across separate periods. If the thresholds were repeatedly changed after viewing all periods, the apparent holdout would no longer be independent. This hypothetical example illustrates test design only; it does not describe a profitable strategy or provide performance statistics.
Conclusion
Backtesting algorithmic trading strategies means evaluating defined rules on historical observations under stated data, timing, cost, and risk assumptions. A backtest can reveal how a process behaves in that simulation, but its conclusions are limited by the design and evidence. Clear documentation, careful interpretation, and separate system testing help prevent historical results from being mistaken for proof of future outcomes. For the wider systematic research process, return to the Algorithmic Trading hub.