At the closing session of a pilot programme, the statement of purpose sitting on the second slide rarely matches, word for word, the purpose recorded in the memorandum that authorised the same pilot six months earlier. The pilot was commissioned to establish whether a new sales channel would reduce customer acquisition cost; the data show no material movement in acquisition cost, while average order value rises appreciably, and the closing deck records the exercise as confirmation of the basket-size thesis. Nobody in the room says anything untrue, no document is altered, and the original authorisation is simply never opened. Viewed across a longer horizon, what emerges is a run of twelve consecutive experiments, twelve of which confirmed something and none of which produced information capable of separating one course of action from another.

The same pattern recurs in other rooms wearing different clothes. In an investment committee session, where realised returns diverge from the thesis that supported the position, the thesis is quietly widened and the holding is reclassified under diversification. In a budget review, a marketing increase justified last year by a growth target that was subsequently missed is reframed this year as an investment in brand awareness. A supplier switch argued on cost grounds, where cost does not fall but lead times shorten, begins to be described as a decision about supply security. The diagnostic signal is identical in each case, and it is procedural rather than moral: the criterion of success is being defined at the meeting where the data are tabled.

The behaviour has a name — HARKing, hypothesising after the results are known, meaning the construction of an explanation once the outcome is visible and its presentation as though it had been specified in advance. The mechanism operates on two levels. The first is cognitive: memory does not retrieve events so much as reconstruct them, and reconstruction is always performed in the light of what is now known, so that the sensation of having expected the result is not a deliberate distortion but the ordinary functioning of recall. The second is organisational: institutions reward coherent narrative because coherence lowers the cost of coordination, and the career difference between reporting that a thesis was disproved and reporting that a thesis was confirmed, while rarely written down anywhere, is real in most structures.

The tendency is functional under identifiable conditions, and no neutralising mechanism can be designed without acknowledging that. In early exploration, where the determining variable is not yet known, looking at the data and forming a new explanation is learning itself; locking an immature line of enquiry to a rigid prior hypothesis suppresses the very signal the exercise was meant to detect. The difficulty lies not in the shortcut but in its persistence after conditions change, because the moment an explanation generated during exploration is reported as though it had been generated during confirmation, the evidential chain supporting the decision is severed. Once exploration and confirmation are compressed into a single sentence, the institution can no longer distinguish what it holds — a finding, or the variation that would be statistically expected whenever enough indicators are tracked.

The least noticed part of the mechanism is the selection effect. A pilot typically generates not one measure but somewhere between fifteen and thirty, and it is an expected rather than an anomalous outcome that several of them will move far enough to appear meaningful on noise alone. Where the hypothesis is chosen after the fact, what gets chosen is usually not the strongest signal but the noise that most resembles one, and when that variation reverts toward the mean in the following period, the organisation reads the reversion as weak execution rather than as evidence that the hypothesis never existed at the outset. Every programme thereby becomes declarable a success; a programme that has lost its falsifiability has, by definition, stopped producing information on which a decision can turn.

The institutional cost surfaces first in the rate of learning. The proportion of initiatives an organisation cancels is the most direct available indicator of what it has learned, and where that proportion approaches zero the correct reading is not that the portfolio is performing but that the portfolio is not being assessed. In practice this appears as the same experiment being re-budgeted for three consecutive years under successive names — a channel trial, then a channel optimisation, then a channel transformation programme — with each round rendered difficult to terminate precisely because the preceding round was recorded as successful. Capital immobilised in an exhausted line is ordinarily more expensive than a mistaken investment, since the cost accrues not in an expense line but in the allocation never made.

The second surface is the diligence table. When a buyer or an investor asks where growth came from, the question being tested is narrower than it appears: whether performance rests on a repeatable mechanism or on an account assembled after the fact. Through diligence, the distinction is drawn by examining whether the attribution of a revenue increase to a particular decision was documented at the time; where the rationale for decisions exists only in post-outcome presentations, the chain of attribution cannot be verified. The valuation consequence is rarely a direct reduction in the multiple. It surfaces instead as a larger earn-out share, a higher escrow ratio, a broader representation and warranty package and an extended founder lock-up — the headline price is preserved while the bearer of the risk changes.

In capital-intensive projects the same tendency meets a harder and earlier surface. Site tests, field measurements and technical pilots conducted before FID feed the assumption set of the financial model, and where the question each input was designed to answer is not on the record, the lender's independent engineer cannot verify the provenance of the assumption and will typically substitute a conservative one. That substitution, expressed through a yield or availability differential that looks minor in isolation, flows into DSCR, compresses debt capacity and enlarges the sponsor equity requirement at closing. The same weakness appears on the contract side: where the performance guarantee is not tied to a specified test protocol, the LD cap becomes a negotiating surface that migrates toward the contractor.

This tendency is not managed through individual awareness, because the deficiency is not one of candour but one of timing in the record. An effective intervention has four components. The first is that at the moment of proposal — before any data exist — a single primary metric, a threshold on that metric, and the result at which the programme will be stopped are committed to writing. The second is that the role defining the success criterion is separated from the role reporting the outcome, a separation that can be established at the level of a signature even in small teams. The third is that secondary findings are labelled as exploration rather than as confirmed results, and are not admitted as inputs to a decision until attached to a fresh test. The fourth is that the closing session opens with a reading of the original authorisation rather than with a summary slide.

On projects BEIREK manages, this operates by keeping the decision record at the point of proposal rather than at the point of approval. Every test — a technical pilot, a supplier trial, a commercial structure trial — enters the scope document with one primary question, one primary measure and a stopping threshold written in advance, while secondary observations are collected in a separate exploration log and carried into the hypothesis pool of the following cycle rather than appended to the result of the current one. The review rhythm is built to match: the closing session opens with the original scope document, and instances in which the outcome contradicts the hypothesis are recorded not as a management failure but as the highest-information output the cycle produced.

The same discipline is applied on the transaction side as assumption lineage. Each material assumption in the financial model carries alongside it the document it was derived from, the date of derivation, and the question it was answering, so that when a counterparty interrogates a number during diligence, the answer is a dated evidential chain rather than a rationale constructed for the occasion. The presence of that chain changes the tone of negotiation measurably, since the buyer's reflex to price risk operates in direct proportion to the number of attributions that cannot be verified.

How much an institution has learned is measured not by how much it has tried but by how many of its trials were constructed in a form capable of disproving their own hypothesis. Whether the rationale for the last twelve decisions can be read from the document that existed when each decision was taken, or only from the presentation delivered in the meeting where the outcome was already known, is a single question that displays, without elaboration, the current threshold of institutional maturity.