In an investment committee session reviewing five years of performance across a development portfolio, the substance of the presentation is almost invariably built from completed assets — realised returns, variance against budget, schedule performance measured against the approved baseline — while projects that were cancelled, that failed during permitting, or that never reached financial close appear as a single line, expressed as a count rather than as a distribution. When the same committee turns, minutes later, to the next allocation decision, the reference set alive in the room consists entirely of those completed assets; the rare but expensive outcomes are present neither as a distribution nor as a recognisable pattern. The same asymmetry is observable in credit committees, in supplier prequalification, and in hiring panels: the decision maker knows the mechanics of the frequent outcome in fine detail and knows the infrequent outcome only by name. This arises from the structure of the sample rather than from inattention, which is exactly why attentiveness alone cannot correct it.
The identical asymmetry surfaces in more measurable form wherever the decision has been delegated to a system. A supplier-default prediction model trained on thousands of purchase orders containing a few dozen defaults can produce an overall accuracy above ninety-nine per cent by never predicting a default at all, and where aggregate accuracy is the only figure reported upward, that model reads as a success. It is not a success; it has memorised the majority class and has never once performed the task it was commissioned to perform, which was to flag the exceptional case. Human judgement and statistical estimation converge here on the same failure for the same reason, and that convergence is itself the diagnostic point: the problem does not live in the software layer, and replacing the model with a better-specified one leaves it untouched. What is at issue is the evidentiary regime under which the rare outcome is permitted to enter the record.
The pattern carries a name — class-imbalance bias, the condition in which one outcome category occupies so overwhelming a share of the observation pool that the minority category becomes neither learnable nor testable. The mechanism operates in two layers. In the first, any learning process, whether it is the parameter calibration of an algorithm or the intuition a manager accumulates across a career, orients itself toward minimising total error, and ignoring the rare class outright is the cheapest available route to a smaller error total. In the second, a feedback loop closes: because the rare outcome is not anticipated, no countermeasure is designed against it; because no countermeasure existed, its eventual occurrence is coded after the fact as a one-off misfortune attributable to circumstances that will not recur; and because it has been coded as singular, it never enters the forward sample as a pattern that could inform the next decision.
There is a range of conditions under which this tendency is genuinely efficient, and designing an intervention without acknowledging that range begins from the wrong premise. Where the outcome distribution is stable, where the cost of the rare event remains within the same order of magnitude as the cost of the ordinary event, and where decisions recur at high frequency, a rule calibrated to the majority class is a sound economising device; examining every low-probability branch individually burns more in review cost than the examination returns. The difficulty lies not in the shortcut but in its persistence after the conditions that justified it have lapsed. In a capital-intensive project where the tail outcome carries a cost an order of magnitude greater than the typical outcome, a frequency-weighted decision rule is no longer efficient — it is simply uncalibrated, and its apparent economy is borrowed against a loss that has not yet been recognised.
The institutional cost appears first in budget and schedule assumptions. Where the duration estimate for a construction programme is derived from the mean duration of completed works, the events that constitute the tail — a permitting objection escalated into litigation, a slipped delivery date on a single-source long-lead item, an interconnection study returned for re-run — never enter the estimating distribution at all, and the contingency line intended to absorb programme risk is sized against the observed body of that distribution rather than against a tail that has never been characterised. This sizing error does not present itself as a discrete entry anywhere in the accounts; it emerges in fragments, through elevated working capital requirements, through unscheduled draw requests, and through exposure to liquidated damages that was priced as remote. A cash flow model can misstate nothing at the level of individual line items and still rest on a wholly incorrect distributional assumption.
The second cost accrues in the contractual and financing layer. Where DSCR tests, reserve account sizing, and covenant thresholds are calibrated against historical variance, and that historical variance excludes the rare class by construction, the resulting thresholds run systematically loose; the cushion visible to the lender is a cushion that has never been tested against the scenario class most likely to consume it. The same logic governs representation and warranty negotiation: where the data room documents only disputes that materialised, leaving exposures that were structurally available but never realised without any documentary trace, warranty scope is negotiated down toward the observed event set, and a tail event surfacing after closing falls within no party's covenant. Escrow proportions indexed to a claims history compiled on the same basis inherit precisely the same blind spot, and inherit it without any party having made an identifiable error.
The third cost sits inside institutional learning itself and is the least visible of the three. Where an organisation retains records only for the projects that survived, its archive becomes progressively less informative over time, because every additional successful file deposited into it increases the weight of the majority class and further depresses the relative visibility of the rare one. A firm with a decade of institutional memory, if that memory consists exclusively of completed work, is not better informed about tail risk than a firm with three years of history; it is merely more confident. At the diligence table this distinction can be read against institutional maturity rather than in its favour, since what a reviewer is looking for is not the volume of accumulated experience but the demonstrated capacity to disaggregate that experience by risk class — a capacity that a survivor-only archive is structurally unable to evidence.
The first component of a structural intervention is to place the rare class under a separate evidentiary regime, so that tail events are not pooled into the same statistic as body events but held in a distinct record under a distinct classification. The second component is asymmetric threshold construction: where the cost of a false negative exceeds the cost of a false positive by an order of magnitude, the decision rule is shifted to reflect that ratio explicitly, and the reasoning behind the shift is committed to writing rather than left as a matter of judgement. The third is the systematic capture of outcomes that did not occur — bids submitted and lost, projects screened and declined, contracts negotiated and abandoned before signature — since these constitute a sample in their own right, and without that sample nothing about the rare class can be learned. The fourth is the relocation of the reporting metric away from aggregate accuracy and onto the capture rate over the rare class.
BEIREK embeds this intervention inside the project management line rather than delivering it as a separate analytical product. In the programmes we run, the decision record opens at the moment of proposal rather than at the moment of approval, so that the basis on which an assumption was selected — the sample it rested on, the alternative that was rejected, the reasoning that connected the two — is committed to writing before the outcome is known, which prevents the outcome from retroactively reshaping the account of how the decision was reached. Alongside this, files that were declined or that lapsed are maintained under the same documentary discipline as files that completed, on the reasoning that the true risk distribution of a portfolio becomes visible only when the unclosed files are counted as observations rather than treated as absences.
Our second line of intervention defines tail scenarios by condition rather than by probability. Instead of estimating how often a given event will occur, we specify which conditions, satisfied jointly, render that event available, and then bind each condition to an observable indicator — the expiry of an objection period in a permitting file, a second deferral of a confirmed delivery date from a single-source supplier, a scope revision request within an interconnection study. What follows when an indicator triggers is not an alert circulated for awareness but a predefined review track operating under a different approval threshold, with a named owner and a defined response window. Pre-mortem sessions form part of this line, and their agenda is not the project's route to success but the mechanism through which its failure would have occurred; the output is not a risk register but a specification of indicators and of the procedure that runs when one of them fires.
Managing class imbalance is, in the end, a question of authority rather than a question of data. The signal that points toward the rare class runs, by definition, against the accumulated experience of the majority, and the individual or function carrying that signal occupies a structurally weak position relative to the decision authority that embodies majority experience; the attenuation of such signals reflects the distribution of authority within the institution rather than any deficit of individual courage. For that reason, unless it has been settled in advance who plays the counter-argument role, under what mandate, and at which stage of the process, no measurement correction will suffice on its own. The question worth putting to an institution is therefore not how prepared it is for the rare event, but whether the person who raised the rare event is taken seriously in the room without first having to be proved right.
