Two backtests can use the same strategy code and produce different results because they do not represent the same market information or execution process. One may use trade-price candles, another bid candles. One may fill every touched limit order, another require evidence of available liquidity. Small differences become important when the strategy's average profit per trade is small.
The productive response is to separate the layers and reconcile them. A chart is a representation of market data. A signal is a decision produced from information available at a specified time. An order is an instruction submitted under defined rules. A fill is the resulting execution. Treating all four as the same event makes discrepancies difficult to diagnose.
Separate platform, data and order routing
The platform displays information and runs trading logic. The data feed supplies prices, quotes and other market events. Order routing carries instructions toward an execution venue or simulation engine. A provider may offer several of these services, but the concepts remain separate. Changing the chart interface does not necessarily change the exchange data, and using the same data does not guarantee identical fills.
For a reproducible experiment, record the instrument identifier, venue, contract, data schema, timestamp convention and fill model. A familiar ticker alone is not enough. A futures root can refer to different expiries, a continuous series or a smaller related contract. Those choices affect prices, monetary value and the history available to the strategy.
Official data documentation is the appropriate place to check whether a schema contains trades, best bid and offer, market depth or aggregated candles. Each answers different questions. OHLCV is useful for many signal calculations, but it cannot reconstruct every quote transition or queue interaction inside a bar.
Understand which price the strategy sees
A last-traded price records an executed trade. A bid is a price currently offered by buyers, and an ask is offered by sellers. A midpoint averages the two. These prices can differ, especially when spreads widen or trades are infrequent. A strategy tested on midpoint prices can look better than a strategy paying the spread to enter and exit.
Consider a hypothetical instrument with bid 100.00 and ask 100.25. Buying immediately at the ask and selling immediately at the bid loses 0.25 price units before commissions. A midpoint-only test that enters and exits at 100.125 would incorrectly show zero. The difference is the spread, not a failure of the trading signal.
Stop triggers depend on the instrument, venue, order type and broker implementation. Do not apply one universal rule that every futures stop uses bid or ask. Check the actual trigger specification and whether the order is held at the exchange, broker or client. The reference that triggers an order and the price at which it fills are separate matters.
Quantify the execution budget
Suppose a hypothetical futures contract is worth $20 per point and has a 0.25-point tick. One tick is therefore $5. A backtest has 200 round trips and $4,000 gross profit before commissions and slippage. Assume a $4 total round-trip commission and initially no slippage.
| Execution assumption | Commission | Slippage cost | Net result |
|---|---|---|---|
| No slippage | $800 | $0 | $3,200 |
| One tick per side | $800 | $2,000 | $1,200 |
| Two ticks per side | $800 | $4,000 | −$800 |
One tick on entry and one on exit costs $10 per round trip, or $2,000 over 200 trades. Two ticks per side cost $20 per round trip. The entry logic did not change, yet the sign of net profit did. This is why the relevant question is not whether the equity curve is attractive before costs, but how much execution error the edge can tolerate.
This simple table assumes every original trade still occurs and only fill prices change. In a complete replay, different fills can change stops, position state and later signals. Use the table as a sensitivity screen, then rerun the strategy when execution assumptions can alter the trading path.
OHLC bars do not establish event order
Imagine a long entry at 100, a stop at 99 and a target at 102. A bar has open 100, high 103, low 98 and close 101. Both stop and target were reachable within the bar. The four summary prices do not establish which was reached first after entry. A backtest that always awards the target is making an optimistic assumption.
Higher-resolution data may resolve the ambiguity, but not always. A one-second bar can still contain both levels. If the necessary event sequence is unavailable, apply a documented conservative rule or present bounds. Do not imply that a deterministic convention reconstructs the actual historical path.
Entry timing creates another common error. If the signal uses the complete close of a one-minute bar, it generally cannot also assume an earlier fill inside that same bar. Specify when the signal becomes available and what is the earliest admissible order event. A visually convincing chart can conceal this timing leak.
A touched limit is not automatically a fill
A limit price appearing in a candle does not establish that your order received execution. Other orders may have priority, the traded quantity may be insufficient or your order may have arrived after the relevant trade. A passive strategy can be particularly sensitive to these assumptions because many of its apparent opportunities occur near brief touches.
A simple model may require the market to trade through the limit by a specified amount. A richer model can use depth and order-book information. Neither is automatically correct for every strategy. State the approximation, test adverse variants and compare with paper or live execution records where available.
Partial fills also affect performance. If a two-contract order fills only one contract, the average exposure and later exit instructions differ from a full-fill assumption. Decide how the strategy handles residual quantity, cancellation and replacement. A simulator should follow that policy rather than quietly pretending the desired quantity always traded.
Timezones and sessions change the candles
A bar boundary is a modeling choice. A daily candle based on UTC midnight is not necessarily the same as an exchange session or broker day. Indicators calculated from those bars can therefore generate different signals. Daylight-saving transitions make a fixed offset especially risky when the strategy depends on local market hours.
Preserve timestamps in a consistent machine representation and retain the relevant named timezone for session logic. Verify whether bar timestamps denote interval start or interval end. A one-minute series labeled 09:30 can refer to different information availability depending on that convention.
Missing bars need context. A period with no trades is different from an exchange closure or a lost data partition. Do not fill every gap with invented volume and a repeated price without documenting the policy. Such filling can alter volatility, holding-time calculations and the apparent number of trading opportunities.
Futures rolls require an explicit policy
A continuous futures chart joins contracts into a research series. It is not itself an executable contract. Back-adjustment can remove visible price gaps while changing historical price levels. That can be useful for some indicators, but cash P&L must still be tied to actual tradable prices and contract specifications.
Record the roll rule and whether the strategy holds through the transition. If positions are closed and reopened, account for both executions and costs. If a signal uses adjusted history but orders trade an unadjusted contract, document the mapping. Comparing two feeds without matching these policies can produce different results even when both feeds are internally correct.
Use historical instrument definitions where contract terms can change. Point value, tick size and symbol mapping belong to the calculation identity. A price difference multiplied by the wrong contract value can create a large P&L discrepancy that has nothing to do with strategy quality.
Reconcile one trade before comparing whole curves
When two reports disagree, choose the first trade whose result diverges. Compare the source events, indicator inputs, signal timestamp, order timestamp, quantity, fill prices and costs. The earliest difference usually explains later differences because strategy state can compound the initial discrepancy.
For a long trade with constant quantity, gross P&L is (exit price − entry price) × point value × quantity. Reverse the price difference for a short. Subtract explicit costs once. If slippage is already reflected in entry and exit prices, do not subtract the same slippage again as a separate expense.
Partial exits require summing individual fills with their quantities. Open equity requires marking remaining exposure. Reconcile the sum of closed-trade net P&L with the change in account balance after deposits, withdrawals and other adjustments. This basic arithmetic check should precede any discussion about why one strategy seems better.
Test transferability rather than visual similarity
A useful validation sequence starts with deterministic historical replay, continues with paper execution using incoming data and then compares any permitted real executions with the model. Record fill shortfall and missed trades instead of judging only the shape of the equity curve. A small sample cannot establish long-term performance, but it can reveal a systematic timing or cost mismatch.
Execution sensitivity varies with strategy structure. Tight-target, high-turnover methods often have less room for spread and slippage than slower strategies. Slower methods still face overnight gaps, financing and roll effects. It is more informative to measure each strategy's cost budget than to label a whole trading style immune to data quality.
- Match the instrument and contract definitions.
- Match price type, session boundaries and bar timestamps.
- Verify signal availability before the assumed fill.
- Stress spread, slippage, missed fills and partial quantity.
- Reconcile individual trades and then the entire equity ledger.
The aim is not to find the feed that makes the backtest look best. It is to identify which information and execution assumptions the result depends on. Once those assumptions are explicit, differences between platforms become testable instead of mysterious.