When a decline in performance reaches the management table, the discussion almost never opens with the question of how many distinct explanations could account for it; within the first ten minutes an explanation is adopted, and the remainder of the session is spent testing that explanation. If a narrowing sales funnel is named a motivation problem, a motivation survey follows the next week, the survey returns low motivation, the diagnosis is treated as confirmed, and the incentive structure is redesigned accordingly. Price positioning, a rival product's feature set, or lost access to the buying committee would each have produced precisely the same survey result, since a team working a narrowing funnel loses motivation whatever the underlying cause happens to be. A test was designed, executed, and honestly reported, and yet the test had no capacity to distinguish between the explanations it was ostensibly weighing.

The same pattern moves more quietly along technical lines, where it is harder to see because the evidence looks physical rather than interpretive. When a plant registers a yield loss, the team examines the component it knows best, takes a measurement on that component, finds a deviation, and replaces it; the possibility that the deviation is itself the downstream signature of a feed fluctuation one stage upstream is never eliminated, because it was never formulated. Yield recovers for a period after the swap, since the replacement part begins operating in the middle of its tolerance band rather than at the edge, and that temporary recovery cements the diagnosis in institutional memory more firmly than any analysis could have done. Six months later, when the same symptom returns, the item under scrutiny is no longer the hypothesis but the supplier's component quality.

The behaviour has a name: congruence bias, the tendency of a decision maker to design tests that will generate results consistent with the preferred hypothesis, while never constructing the test that would separate that hypothesis from its rivals. It belongs to the same family as confirmation bias but operates through a different mechanism. Confirmation bias concerns the selective reading of evidence already in hand, whereas congruence bias sits one step earlier, in how the evidence is generated — that is, in the architecture of the test itself. Nobody distorts the data; the test is run in good faith and the result is reported in good faith, and what escapes notice is that the test was never discriminating. The danger lies precisely in that good faith, because from the outside the process presents every appearance of a properly conducted analysis.

Underneath the mechanism sits a saving in computational effort, and in most circumstances that saving is entirely rational. Formulating rival hypotheses, defining a discriminating observation for each, and building the measurement arrangement capable of collecting those observations costs several times what testing a single hypothesis costs. For decisions that repeat, carry low variance, and can be reversed, the additional expenditure does not earn its keep; a wrong diagnosis is corrected on the next cycle and learning is cheap. The difficulty lies not in the shortcut itself but in the shortcut persisting after the conditions that justified it have changed. Once a decision becomes irreversible, once the capital commitment scales, or once the diagnosis is embedded in the structure of a contract, the same reflex continues to operate and learning has ceased to be cheap.

Organisational structure feeds the tendency more powerfully than individual disposition does. A hypothesis acquires weight in proportion to the seniority of whoever voices it first, and proposing a competing account of the same evidence tends to be received not as an analytical contribution but as an implicit objection to a colleague. The brief handed to the analytical team is usually written in the same direction: a request phrased as validate this hypothesis, or model this scenario, produces by construction a single-hypothesis study, and the more rigorous the team's output, the more solid the diagnosis appears. Rigour functions here as an amplifier rather than a corrective, since executing a poorly designed test with great care does nothing except narrow the confidence interval around a wrong conclusion, which is exactly the outcome an approving committee will read as strength.

The institutional cost is booked not against the first intervention but across the second and third rounds. Because the initial investment built on a wrong diagnosis — a redesigned incentive plan, an additional equipment line, an ERP module, a new regional office — suppresses the symptom for a period, it is recorded as successful and becomes a reference point in the following budget cycle. When the symptom returns, what comes under question is the adequacy of the implementation rather than the diagnosis that justified it, so a second layer is constructed on the same faulty premise and total commitment reaches several times what a single, correctly targeted intervention would have required. The accumulation is most visible not in the fixed asset line but in the operating expenditure attached to it, since a structure once established continues to carry maintenance, headcount, and depreciation long after its premise has been abandoned.

At the diligence table the pattern surfaces through one specific question: for each of the three largest operational decisions of the past three years, which competing explanation was considered against the adopted diagnosis, and which observation eliminated it. The number of companies able to produce a written answer is limited even among those with disciplined documentation, since the decision record typically captures the option selected and the rationale supporting it, while omitting the explanations discarded and the criterion by which they were discarded. What the buyer infers is that decision quality derives from the judgement of particular individuals rather than from an institutional method, and that inference is priced under the heading of key-person dependency. In the valuation discussion it rarely appears as a general reduction in the multiple; it appears as a longer observation window in the earn-out structure, or as a payment tranche conditioned on the retention of named personnel.

The same weakness presents a harder edge in technical diligence, where it can be traced through documents rather than inferred from behaviour. Where an asset's performance history contains a recurring fault pattern, the reviewing party's interest lies not in the fault itself but in which hypotheses the corrective action record evaluated after each occurrence; a record showing a single root cause and a single intervention indicates, structurally, an elevated probability that the same fault will recur once the warranty period has expired. That assessment then finds expression across several headings simultaneously — the insurance premium, the sizing of the maintenance reserve, the breadth of representations and warranties, and the list of conditions precedent to closing. What the seller loses at that point is not a technical score but negotiating ground, and the loss compounds because each heading is argued separately.

This tendency is not manageable through individual awareness, since a person cannot, by definition, observe from inside a test design that the design lacks discriminating power. What neutralises it is institutional architecture, and in practice that architecture separates into three components. The first is keeping the decision record at the moment of proposal rather than the moment of approval, and making the naming of at least one rival hypothesis a formal condition of the entry. The second is writing down the discriminating observation for each hypothesis in advance — answering, before any data is collected, the question of which measurement result would eliminate this explanation. The third is separating the person who proposes the diagnosis from the person who designs the test, so that the design is constructed to discriminate rather than to protect an account somebody already owns. Operating together, these three leave the team's speed intact while rendering its diagnoses testable.

On projects under BEIREK's management this architecture operates as a fixed field in the decision log: for every material technical or commercial diagnosis, the entry carries the preferred explanation, at least one competing explanation alongside it, and the observation that would separate the two, and a proposal reaching the agenda with that field left blank is returned rather than discussed. The requirement is not a formality but the leading edge of scope and cost control, since in our experience a meaningful share of change orders originates not in poor execution but in a root cause that was never discriminated during the first round, and which is then corrected in the second round through a far more expensive contractual amendment. The same discipline governs the assumption register maintained during development, where each critical assumption is paired with the observation capable of invalidating it, and that observation is tied to a specific milestone in the project schedule.

The second mechanism is defining the counter-argument function as a permanent and impersonal role rather than a temperament. In investment committee sessions and project steering meetings, the responsibility for presenting explanations that compete with the proposed diagnosis is assigned on rotation; the output of that role is not an objection but a proposed discriminating test, and it is minuted under its own heading rather than absorbed into general discussion. Rotation is essential, because a permanent dissenter loses organisational effect over time as the substance of the argument is gradually attributed to the personality making it, and the weight of the contribution declines regardless of its quality. Where the rhythm holds, generating a rival hypothesis stops being a career risk and becomes a routine procedural step — and that shift, rather than any analytical technique, is the actual lever governing this tendency.

What demonstrates the robustness of a decision is not the volume of evidence supporting it but the number of alternative explanations that evidence eliminates, and where the two are conflated an organisation carries its greatest exposure precisely in the analyses into which it has poured the most effort. One question, cheap to ask and uncomfortable to answer, belongs on the management table before any material commitment is approved: if this diagnosis were wrong, which of the observations already in hand would look different.