A reproducible backtest is one that another run can reconstruct from its saved inputs, not merely one that produces an attractive equity chart. You need to know which code ran, which parameter values it used, what data it saw, how orders were filled and which costs were charged. If any of those are ambiguous, a later report can disagree without either screen explaining why.
The practical goal is a traceable chain from source data to signals, fills, trades, equity and summary metrics. This article follows that chain with a small worked example. It also explains how input settings such as A, B and C can keep experiments understandable without treating every analysis phase as a new strategy.
Separate code from its input values
The code defines the trading logic. Inputs choose values within that logic, such as a lookback length, threshold or permission to short. Changing lookback from 20 to 25 need not create a new code version, but it does create a distinct parameter setting. Changing the entry formula itself requires a new code version even when the input names remain identical.
For example, setting A might be lookback 20 and threshold 0.15. Setting B might use 25 and 0.15. Setting C might use 25 and 0.20. A compact letter makes the settings easy to select. Their identity still needs the exact typed values and code version. The letter is a label, not the mathematical definition of the strategy.
A report belongs to its precise setting. Validation of B adds evidence about B on another period. It does not become an unrelated candidate whose parameters must be guessed later. If an input changes, generate or reuse the correctly matching setting rather than attaching B's previous profits to the new values.
Record the whole calculation context
Matching inputs alone is insufficient for reuse. Instrument, bar interval, date boundaries, contract mapping, data version, capital, execution model and costs can all change a result. A run over January through March is not interchangeable with a run over January through May, even if both use setting B.
Use explicit period conventions. An interval starting January 1 and ending April 1 can be represented as start-inclusive and end-exclusive, which avoids double-counting a timestamp shared with the next period. The important point is consistency. Store timezone-aware boundaries and define how trades crossing a boundary are handled.
Randomized procedures need their random seed and algorithm version. Deterministic backtests should not depend on unrelated application state, the last report opened or the order in which the user visited screens. If cached results are reused, their key should describe the calculation rather than a convenient display label.
Make information availability explicit
A strategy cannot use a bar's final high, low or close before that bar has completed. If a signal is calculated from the close, define the earliest subsequent event at which an order may execute. Filling at the same close can be a legitimate convention only under an explicitly modeled order process, not an automatic assumption that a completed bar was known in advance.
Indicators need warm-up data. A 20-bar moving average cannot be computed from five observations without changing its definition. Load sufficient earlier history for calculation, then begin performance measurement at the intended boundary. Do not silently include warm-up trades in the measured sample.
External features have their own availability timestamps. A macro observation dated January may not be published until February and may be revised later. A weekly positioning report describes an earlier date than its release. Join such features according to when the strategy could know them, not simply their observation labels.
Build a trade ledger you can audit
For every trade, retain direction, entry and exit times, quantities, actual modeled fill prices, instrument value and each cost component. For partial exits, retain the individual fills instead of compressing them into an average that loses quantity information. The ledger should explain the summary without requiring the chart.
Consider a hypothetical contract worth $20 per price point. The following trades use fixed quantities and explicit round-trip costs. Slippage, if any, is already reflected in the listed fill prices, so it must not be subtracted again.
| Trade | Direction and quantity | Entry to exit | Gross P&L | Costs | Net P&L |
|---|---|---|---|---|---|
| 1 | Long × 1 | 100 to 110 | $200 | $10 | $190 |
| 2 | Short × 2 | 110 to 106 | $160 | $20 | $140 |
| 3 | Long × 1 | 105 to 102 | −$60 | $10 | −$70 |
| Total | 3 trades | Closed positions | $300 | $40 | $260 |
Long gross profit is exit minus entry, multiplied by point value and quantity. For a short, entry minus exit gives the correct sign. Trade two therefore earns (110 − 106) × $20 × 2 = $160 before costs. The three net results sum to $260. This simple test catches sign, quantity and double-cost errors before more complicated analytics are introduced.
Reconcile equity and metrics
With starting capital of $100,000 and no external cash flows, the closed-trade equity sequence is $100,190, $100,330 and $100,260. The final value is starting capital plus $260. The largest closed-equity drawdown is $70 from the second-trade peak to the final value. It is not proof that intratrade drawdown was only $70.
For this net trade ledger, win rate is two divided by three, or 66.67%. Net-based profit factor is $330 of winning net outcomes divided by $70 of losing net outcomes, or approximately 4.714. Average trade is $260 / 3 = $86.67. A report using gross profit factor would produce a different number and should label the convention.
Do not estimate a meaningful annualized Sharpe ratio from these three trade outcomes without a time model. A Sharpe ratio normally uses returns sampled over defined intervals, with assumptions about annualization and dependence. A missing statistic can be more honest than a number computed from incompatible units.
Separate closed equity from marked equity
Closed equity changes when trades realize profit or loss. Marked equity also includes open positions. These curves answer different questions. A strategy can close every trade profitably while suffering large temporary losses. For account-level risk, margin or prop-limit analysis, those temporary losses can matter.
Choose a consistent mark price and timestamp frequency. If only OHLC bars are available, the exact intrabar path may be unknown. Do not make the open-equity curve appear more precise than the data supports. A conservative approximation should be identified as a model rather than described as a reconstructed tick path.
When comparing phases, state whether starting equity resets at each boundary or continues cumulatively. A validation curve normalized to $100,000 is useful for standalone comparison. Concatenating that reset curve directly after development without adjusting the level creates a false jump. Preserve the distinction between segment performance and one continuous portfolio path.
Use chronological data splits
Development data support design and optimization. Validation checks choices on later or otherwise separated data. A final holdout is reserved for a locked decision. The split does not make a strategy good, but it helps distinguish designing a rule from evaluating a rule that was already chosen.
Parameters, filters and objectives should be selected without reading the final holdout. If you inspect the holdout and change the strategy, that data has influenced development. You can still rerun it, but it is no longer fresh independent evidence. Store that history rather than pretending a new code label makes previously inspected observations unseen.
Boundary handling needs care when trades overlap periods. You can assign trades by entry, close positions at a boundary, or use another declared convention. Each creates a different experiment. Ensure metrics and charts use the same convention, and do not accidentally count one trade in two reported segments.
Test the engine before trusting a complex strategy
Use small artificial price sequences with known outcomes. Test one long, one short, a gap through a stop, a target and stop inside the same bar, a partial exit and a trade crossing a session boundary. Verify exactly which event wins when conditions coincide. These fixtures expose assumptions that a large historical backtest can hide.
Also test invariants. Net profit should reconcile with the ledger. Closed quantity should not exceed opened quantity. A flat account should have no unrealized exposure. Repeated runs with identical inputs should agree. Changing a display filter should not mutate saved results. These checks protect the calculation chain rather than one favorable example.
For cached results, test both reuse and invalidation. Identical code, inputs and context can reuse a verified artifact. A change in commission, data content or engine version must not retrieve an incompatible result merely because the strategy name is unchanged. Correct reuse improves speed without mixing evidence.
Compare settings without losing identity
Once the base run is reconciled, compare A, B and C across the same data and cost assumptions. Preserve each setting's separate report. If B is selected for validation, its validation result should refer to B's exact inputs and code, not whichever setting happens to be active on the screen afterward.
Keep analysis-specific settings separate too. Monte Carlo simulation count does not change the historical trade ledger. An alternative exit rule does. A macro filter explored after the run is not automatically a new executable strategy until it is implemented and replayed with the necessary position-state changes.
A practical handoff includes the saved configuration, trade ledger, equity series, metric definitions and software version. The report then becomes a view of reproducible evidence rather than the only place where the evidence exists.
A reliable first milestone
Before optimizing hundreds of settings, make one setting fully explainable. Reconcile every trade, confirm the period and costs, verify the timing of signals and rerun it from the saved configuration. That small milestone provides a dependable reference for every later phase. Fast optimization on top of an ambiguous base run only produces more ambiguous results.