---
title: "The Garden of Forking Paths: How Unrecorded Analytical Choices Manufacture Invisible Multiple Testing"
description: "The garden of forking paths arises when small definitional, sampling and threshold choices are made after looking at the data, so that a single analysis is presented while many alternative analyses have effectively been explored. Its institutional signature is the base case that clears committee but erodes in operation; the neutralising mechanism is fixing the analytical protocol in writing before any result is visible."
url: https://www.beirek.com/en/blog/garden-of-forking-paths
canonical: https://www.beirek.com/en/blog/garden-of-forking-paths
published: 2025-05-02
modified: 2025-05-02
category: "Organisational Psychology"
category_url: https://www.beirek.com/en/blog/category/organisational-psychology
language: en-US
reading_time_minutes: 7
publisher: BEIREK LLC
publisher_url: https://www.beirek.com
license: "© BEIREK LLC — citation with attribution and link permitted"
keywords: ["garden of forking paths","analytical protocol","base case reforecast","normalisation adjustments","due diligence findings"]
topics: ["Decision architecture in investment committees","Model governance and assumption logging","Valuation adjustments and transaction structuring"]
alternate_language_url: https://www.beirek.com/tr/blog/garden-of-forking-paths
---

# The Garden of Forking Paths: How Unrecorded Analytical Choices Manufacture Invisible Multiple Testing

> **In short:** The garden of forking paths arises when small definitional, sampling and threshold choices are made after looking at the data, so that a single analysis is presented while many alternative analyses have effectively been explored. Its institutional signature is the base case that clears committee but erodes in operation; the neutralising mechanism is fixing the analytical protocol in writing before any result is visible.

*A single figure placed before an investment committee is, more often than not, the composite of dozens of small analytical choices that nobody logged as decisions. Made after the data has been examined, those choices produce a sequence of tests that appears never to have been run, and the model's true confidence interval is materially wider than the one the committee sees.*

---

What arrives on the table at an investment committee is a single number: the average debt service coverage ratio the project will produce across the contract term, or the normalised operating margin of the target company. Standing behind that number are dozens of choices, not one of which appears in the minutes as a decision — which quarter serves as the base, which one-off item is admitted into the normalisation, which spot rate governs the currency translation, whether customer concentration is measured across the top five names or the top ten, whether an anomalous month remains in the sample. Each of these looks indisputable at the moment it is made; in the analyst's mind they are not decisions at all but the correct execution of the work. The output that reaches the committee is therefore a single branch of a tree that was never drawn, and the question of how far that branch sits from the others goes unanswered because it goes unasked.

The pattern is hardly confined to the finance desk. A commercial team measuring the effect of a price change on conversion decides which weeks to exclude for promotional distortion only after the series has been plotted; supplier performance in a manufacturing plant acquires meaning only once someone has settled which delays count as logistics and which as production; a pilot is deemed successful once the success threshold has been read off whichever metric proved most legible after the pilot closed. In every instance the organisation believes it has taken one measurement, and it carries the result of that measurement forward as a single fact. The measurement itself, however, has been shaped by a sequence of small structuring choices taken in view of the outcome.

The name for this mechanism is the garden of forking paths — the condition in which every choice of definition, sample, segmentation and threshold is made after inspecting the data, so that although one analysis has been executed, a large number of alternative analyses have implicitly been explored. The analyst does not try five segment definitions in sequence and keep the best; rather, what the first chart reveals determines which segment definition appears natural, and that definition is adopted. Branches never taken are never computed and therefore leave no record, yet the fact that the result landed on this particular branch is conditioned on the data itself, which in statistical terms is indistinguishable from having run multiple tests. The confidence interval presented carries only the uncertainty of the final calculation; it carries none of the uncertainty generated by the selection process that produced it.

This is structurally distinct from deliberate filtering — from p-hacking, the repetition of an analysis until it yields the desired answer — and that distinction is precisely what makes it more persistent in a corporate setting. Deliberate filtering is a question of integrity, and once detected it can be named. Forking paths, by contrast, emerge within the ordinary working rhythm of a well-intentioned, experienced and technically competent analyst, with no rule broken anywhere. Internal audit, checking the model line by line, finds no error, because there is none; what is missing is not a correct figure but a map of the roads not taken. The tendency is consequently unmanageable through checklists and manageable only by altering the temporal order of the process itself.

There is a functional side to the tendency, and designing an intervention without acknowledging it slows analytical capacity to no purpose. Sector judgement consists precisely in knowing which data carry information and which carry noise; when an engineer who knows that a three-day stoppage on a production line was maintenance-driven excludes those days, the data are corrected rather than corrupted. Nor is the cost of computing every defensible branch zero — opening every combination of definitions can turn a week of work into a month of it. The problem lies not in the shortcut but in the absence of a record wherever the direction of the shortcut is correlated with the result; judgement is legitimate, whereas unrecorded judgement cannot be verified.

The institutional cost does not present itself as a modelling error; it presents itself in the calendar and in the cash flow. Where the base case underpinning an investment decision cleared the hurdle rate comfortably in committee yet is revised downward systematically across the three or four quarters following first draw, the probable explanation is not an isolated forecasting failure but a chain of choices that ran through the optimistic branches at each fork. Every choice is individually defensible; accumulated in one direction, they place the result at the upper end of a defensible distribution. The observable trace of that accumulation is the non-randomness of the revisions: where downward corrections arrive markedly more often than upward ones, the matter under examination is not forecasting skill but analytical architecture.

On the transaction side the cost crystallises as a diligence finding and, from there, as price mechanics. Once buy-side advisers open up the adjustments through which normalised EBITDA was derived, and observe that each adjustment is separately defensible while the aggregate runs in a single direction, the discussion migrates from any individual line item to the method as a whole. The typical consequence is less a headline price reduction than a migration of risk into the structure: greater weight on the earn-out, a higher escrow percentage, a widened representation and warranty package under the financial information heading, or an independent recomputation imposed as a condition precedent. For the seller this means that identical performance converts into cash later and on more conditional terms.

On the operating side the cost collects in pilots that fail to scale. A pilot appears successful because its success criterion was selected afterwards; the rollout budget is approved against that appearance; and when the effect does not reproduce in the field, the diagnosis is sought in execution discipline rather than in measurement design. On the lending side the same mechanism surfaces in covenant calibration: where a DSCR threshold is built upon a projection series whose definition was shaped, along the way, by the most favourable branch, the covenant is touched earlier than expected in the first stress period, and the ensuing waiver negotiation occupies the project's management agenda for months.

What neutralises the tendency is not individual vigilance but four components that reorder the process. The first is committing the analytical protocol to writing at the moment the question is posed and before the data are examined: which period, which segmentation, which normalisation rule, which stated grounds for exclusion. The second is a branching map — listing, while the protocol is drafted, the alternative definitions considered defensible, and computing the result not on one branch but across every combination, so that it is presented as a distribution; the basis of decision then becomes the width of that distribution and the proportion of branches clearing the threshold, rather than a single point. The third is holding a portion of the data set closed until the analytical design is complete, keeping the validation window separate. The fourth is fixing the decision threshold before any result is visible, since a threshold that can be moved afterwards strips the analysis of informational value.

The intervention BEIREK operates in capital-intensive projects follows that sequence. An assumption protocol is written and frozen before the base case is constructed; each assumption is entered in an assumption log with its owner, its rationale and the conditions under which it may be changed, so that a later revision appears not as a model update but as a named, dated change of choice. Sensitivity axes are set by the committee that will take the decision rather than by the team performing the analysis, which structurally reduces the probability that the most favourable branch is selected, since the combinations to be tested are fixed while the result is still unknown. In transaction and lending workflows, the directional distribution of normalisation adjustments is additionally maintained as a separate schedule, so that instances in which every adjustment runs the same way are opened and reasoned through internally before the counterparty raises them.

Whether this rhythm has taken hold inside an organisation is measured simply, and the measurement requires roughly a year of observation: has the direction of reforecasts balanced out over time, or does it still fall predominantly one way? In the second case, what warrants examination is not the forecasting ability of the analysts but the point in time at which, and the audience before whom, the organisation makes its analytical decisions — because the reliability of a number depends as much on whether the road leading to it was chosen before or after the result became visible as it does on the arithmetic of the final calculation.

## Key Points

- Any presented result is one branch of a tree built from choices about definition, period, segmentation and threshold; where the tree is never drawn, the robustness of the branch cannot be assessed.
- The mechanism differs from deliberate selection in that the analyst genuinely runs one analysis, yet advances through each fork by consulting the data, so multiple-testing exposure accumulates without leaving an audit trail.
- The institutional cost surfaces not as a modelling error but as a base case that erodes between signing and first draw, followed by a reforecast cycle whose corrections run predominantly in one direction.
- Writing the analytical protocol at the moment the question is posed makes any later change of definition visible and therefore contestable rather than invisible and therefore unchallenged.
- A branching map requires the result to be recomputed across every defensible combination of definitions, so that the decision rests on the width of the resulting distribution rather than on a single point estimate.

## Questions

### How does the garden of forking paths differ from p-hacking?

P-hacking is the deliberate repetition of an analysis until the desired result emerges, and it is fundamentally a question of integrity. The garden of forking paths arises while a single analysis is being run: as the analyst examines the data, the choice of definition, period or threshold that appears natural is itself conditioned on what the data show. The outcome carries the statistical signature of multiple testing, yet because no rule is broken, it leaves no audit trail.

### How can one tell whether a model has been affected by this bias?

The most reliable indicator is not the verification of any single calculation but the directional distribution of revisions. Where corrections to the base case fall predominantly one way over time — usually downward — the question to raise concerns analytical architecture rather than forecasting skill. A second indicator is the sign of normalisation adjustments: where each adjustment is defensible in isolation while the aggregate pushes in one direction, the chain of choices was plausibly selected in relation to the result.

### Does fixing the protocol in advance not disable the analyst's sector judgement?

No; what is fixed is the timing of judgement, not judgement itself. An engineer who knows a stoppage was maintenance-driven and excludes those days corrects the data rather than distorting them. The protocol simply requires that this decision be taken before the result is visible and recorded with its rationale, which makes the same judgement contestable afterwards and testable against alternative definitions. Judgement left unrecorded cannot be verified and is therefore discounted by the counterparty.

### How does this bias affect transaction pricing?

Its effect usually appears not as a direct price reduction but as risk pushed into the structure. Once it becomes visible that normalised figures rest on adjustments running consistently in one direction, the typical buy-side response is to increase the weight of the earn-out, raise the escrow percentage, widen representations and warranties under the financial information heading, or require an independent recomputation as a condition precedent. For the seller, the consequence is that identical performance converts into cash later and on more conditional terms.

---

Source: https://www.beirek.com/en/blog/garden-of-forking-paths
Publisher: BEIREK LLC — https://www.beirek.com
