The best development result is not always the most repeatable strategy. Sometimes it is simply the specification most closely adapted to one sample. A useful way to investigate this is to repeat the act of choosing a winner on different parts of the same historical dataset and then inspect that winner elsewhere. You are evaluating the selection process, not merely admiring its final output.

Suppose an optimizer compares input sets A, B, C and D. A dominates an early trending segment. C dominates a later rebound. B earns modest but steady results. Depending on which segment you use for selection, either A or C can appear exceptional. If that winner consistently loses its relative rank on the complementary sample, the selection rule deserves scrutiny.

What PBO asks

The Probability of Backtest Overfitting framework of Bailey, Borwein, López de Prado and Zhu uses combinatorially symmetric cross-validation to study the out-of-sample ranks of in-sample winners. Historical observations are divided into blocks, combinations of blocks form selection samples, and the complements form evaluation samples. A rank below the middle indicates that the selected winner underperformed the typical available alternative in that complementary evaluation. The reported frequency is conditional on the candidate set, historical sample and selection method, not a universal probability of future failure.

That distinction matters. PBO does not answer “Will this exact strategy lose money tomorrow?” A selected strategy can be profitable yet rank poorly compared with its alternatives. It can also rank well while every candidate loses. Relative selection quality and absolute profitability are separate questions, so inspect both.

The data structure you need

Start with a matrix whose rows represent the same observation dates and whose columns represent different specifications. Each cell should contain a comparable net return or another carefully defined performance observation. Missing trades do not automatically mean missing data. If a strategy was flat on a valid trading day, a zero return may be the correct observation. If the source data were unavailable, replacing the missing result with zero would hide a data problem.

Every column needs the same economic definition. Do not compare one candidate at one contract with another at five contracts using total profit and call the result a pure parameter test. Decide whether you are studying signals, risk sizing or complete executable strategies. Then make the interpretation match that decision.

A second requirement is the full candidate set. Keeping only your five favorite backtests after examining hundreds understates the search. The procedure should see the alternatives that genuinely competed during development, including unattractive ones. A neat spreadsheet of survivors is not a faithful research history.

A small example you can reproduce

The table below is a deliberately constructed teaching dataset. Each number is a block profit in hundreds of dollars under equal capital and exposure assumptions. We use average block profit as the selection score to make the arithmetic visible. This is not a production Sharpe-based study, and six coarse blocks are not sufficient evidence for trading decisions.

BlockABCD
181−42
292−20
3−3192
4−5280
521−14
6−1211

Select on blocks 1, 2 and 3. A averages 14/3, or 4.667 units, and wins. On blocks 4, 5 and 6, A averages −4/3, or −1.333 units. The other averages are 1.667 for B, 2.667 for C and 1.667 for D. A is last in the evaluation sample, despite being first in the selection sample.

Reverse the halves. C wins selection on blocks 4, 5 and 6 with 2.667 units. Its complementary average is 1.000, below A at 4.667, B at 1.333 and D at 1.333. This reversal is not a second independent experiment. It is a second view of the same small historical dataset.

The same candidates across six historical blocksExact teaching matrix from the article. Each selection uses three blocks and evaluates the winner on the complementary three, with no additional observations introduced.The same candidates across six historical blocks89-3-52-1121212-4-298-11202041123456ABCDPerformance units per blockHistorical blockCandidate
Exact teaching matrix from the article. Each selection uses three blocks and evaluates the winner on the complementary three, with no additional observations introduced.

From ranks to a summary

With four candidates, assign evaluation ranks from one for worst to four for best. Use average ranks for tied evaluation scores. Divide the rank by five, which is the number of candidates plus one. This keeps the relative rank away from the endpoints zero and one. Its logit is the logarithm of relative rank divided by one minus relative rank.

relative rank = evaluation rank / (candidate count + 1)
rank logit = log(relative rank / (1 − relative rank))
below-median event = rank logit < 0

For A in the first split, relative rank is 1/5, or 0.20. The logit is approximately −1.386. The negative sign records a below-middle evaluation rank. It does not mean a loss of 138.6%, a return forecast or a leverage recommendation.

There are 20 ways to choose three blocks from six. If we evaluate all 20 selections in this toy dataset, use alphabetical order to break tied selection scores, average tied evaluation ranks and count only strictly negative logits, 19 selections produce a below-middle evaluation rank. One lands exactly in the middle. The resulting diagnostic frequency is 95%. Its precision should not be confused with certainty about the future.

Selection blocksWinnerEvaluation meanRelative rank
1, 2, 3A−1.3330.20
1, 3, 5D0.3330.20
1, 3, 6C1.6670.50
2, 4, 6C1.3330.40
4, 5, 6C1.0000.20
Where the selected winner ranks after selectionAlphabetical selection ties and average evaluation ties. Nineteen logits are negative and one is zero, giving the stated 19/20 diagnostic frequency. These are not independent future trials. Below the evaluation middle: -1.386, -1.386, -1.386, -1.386, -1.386, -1.386, -1.386, -1.386, -1.386, -1.386, -1.386, -1.386, -0.847, -0.405, -1.386, -1.386, -1.386, -0.847, -1.386. Exactly at the middle: 0Where the selected winner ranks after selectionBelow the evaluation middleExactly at the middle-2-1.5-1-0.500.515101520Split index (lexicographic selection blocks)Evaluation rank logit
Alphabetical selection ties and average evaluation ties. Nineteen logits are negative and one is zero, giving the stated 19/20 diagnostic frequency. These are not independent future trials.

Read the failure pattern, not just the percentage

The useful next question is why the selected winners reverse rank. In our example, A and C concentrate their strongest profits in different blocks. The selection rule repeatedly follows whichever concentration happens to be included. B is steady, but its modest returns rarely win the selection contest. The optimization objective rewards historical peaks rather than cross-block consistency.

That does not automatically mean you should replace the winner with B. The dataset was constructed to illustrate a mechanism. In real research, you would investigate exposure, regime dependence, transaction costs and the possibility that different candidates have valid but different economic roles. Selecting a “stable” candidate after looking at every diagnostic is itself another selection decision that needs fresh evaluation.

Inspect the distribution of degradation as well. A small ranking reversal among very similar candidates is economically different from a collapse from strong profit to a large loss. Record changes in average return, drawdown, turnover and performance relative to a simple benchmark. A single frequency cannot describe all of those dimensions.

Selection performance versus complementary evaluationEach marker represents a selected winner from one of 20 dependent splits. Coincident markers overlap. The diagonal means no score change. This is score deterioration, not the rank-based PBO statistic. Selected candidate A: -1.333, -0.667, -3, -2, 0.333, 0.667, 1.333, 0. Selected candidate B: 1.333. Selected candidate C: -0.667, 1.667, -1.333, 1, 1.333, -1.667, -2.333, 0.667, 1. Selected candidate D: 0.333, 1. Unchanged score: -4, 8Selection performance versus complementary evaluationSelected candidate ASelected candidate BSelected candidate CSelected candidate DUnchanged score-5-2.502.557.510-4-202468Selection mean (performance units)Evaluation mean (performance units)
Each marker represents a selected winner from one of 20 dependent splits. Coincident markers overlap. The diagonal means no score change. This is score deterioration, not the rank-based PBO statistic.

Block design and leakage

Blocks should preserve the intended unit of time and enough local structure for the score to make sense. Extremely small blocks create noisy measurements and can cut through positions. Extremely large blocks leave few distinct arrangements. Neither choice has a universal optimal value. Examine a predeclared range and explain which time dependence the design is intended to preserve.

Boundary handling is critical when trades overlap or signals use long histories. A trade's outcome must not be partly visible in the selection data and partly reused as evaluation evidence without a documented rule. Removing overlapping labels or adding a separation gap can help, but a gap chosen from convenience rather than the information horizon may be inadequate.

Combinatorial splits are also not a literal simulation of live chronological retraining. Some arrangements select using later blocks and evaluate earlier blocks. That is part of this diagnostic construction, not permission to deploy a model with future information. A separate forward-only walk-forward exercise answers the operational question of what could have been chosen at each historical date.

A practical implementation checklist

  1. Freeze the candidate universe and the selection objective.
  2. Reconcile each candidate's net returns on the common calendar.
  3. Define blocks, boundary rules and tie handling before inspecting results.
  4. For every split, select using only the selection score.
  5. Evaluate every candidate on the complement and rank the selected winner.
  6. Store each split, winner, rank and absolute performance change.
  7. Compare the diagnostic with a genuinely chronological validation process.

Use immutable identifiers for code, input set, data snapshot and costs. The same letter B may be a convenient display label, but it should resolve to one precise specification. If a parameter changes, do not overwrite the old column while keeping its old returns. Otherwise the matrix no longer describes any real set of experiments.

Test your implementation on tiny matrices where every result can be inspected by hand. Verify that a consistently superior column stays near the top, that reversed performance patterns produce poor complementary ranks and that tie rules are deterministic. Only then scale to a large optimization archive.

How to act on the result

A high overfitting diagnostic invites a simpler hypothesis, a smaller justified search and better separation between selection and evaluation. It does not invite repeatedly changing the block design until the percentage falls. That would optimize the diagnostic itself. Write down which changes are economically motivated, then test them with information not used to choose those changes.

A low value is encouraging only within the tested setup. It cannot detect every data error, guarantee realistic fills or anticipate a new market regime. It also does not excuse an unrecorded history of discarded ideas. Treat PBO as one view of research repeatability alongside execution audits, cost sensitivity and forward evidence.

The central lesson is practical: save the alternatives, not just the winner. Once the full experiment history is available, you can ask whether your method of choosing a strategy survives changes in the sample. Without that history, a beautiful equity curve tells you much less than it appears to.