When a board presentation reports that the risk early-warning system operated at a ninety-seven percent hit rate over the preceding quarter, the reaction around the table is typically approval; the system is working, its budget is renewed, and the team's performance review closes favourably. In the same meeting, the fact that none of the three material incidents that actually occurred during that quarter had been flagged in advance sits under a separate agenda item, often in the annexes to the operations report. Both pieces of information reach the same room, at the same hour, in front of the same people, yet the logical relationship between them is never established, since the first has been classified as a performance metric and the second as an incident log. This separation is not a lapse of attention; to the extent that the reporting architecture itself files the two items in different sections, the work of connecting them is quietly delegated to the individual effort of whoever happens to be reading closely.
The same pattern recurs across credit risk scoring, supplier quality screening, internal audit sampling, fraud detection, occupational health and safety whistleblowing channels, and any classification model built on machine learning. The shared condition is straightforward: the event being measured is rare, while the state in which nothing happens is dominant. Under that condition, a system that flags nothing whatsoever — one that classifies every case as clean — will report ninety-seven percent accuracy where the underlying incidence of the event is three percent. The metric is arithmetically correct, the calculation withstands audit, and the number survives review without objection; what it carries, however, is not a statement about the capability of the system but a restatement of how uncommon the event happens to be.
The structure at work here is the accuracy paradox — the condition under which a rising overall hit rate can coincide with a falling capacity to find the very thing the system exists to find. The mechanism operates on two layers. The first is arithmetic: under an imbalanced distribution, selecting the majority class in every instance produces the highest score when performance is viewed through a single metric, so any system operating under optimisation pressure will drift naturally in that direction. The second layer is organisational and, in practice, more determinative: once a metric is attached to performance appraisal, budget allocation or a management report, raising that metric becomes rational behaviour for the team responsible for it, and in a rare-event domain the cheapest available route to raising it is to lift the threshold, which is to say to flag less.
Reading this tendency as a defect would misstate the case. An overall accuracy figure is a highly functional summariser under conditions in which the classes are balanced and the costs of the two error types sit close to one another; its reduction to a single number is precisely what permits comparison, time-series tracking and benchmarking across business units that otherwise share no common language. The difficulty lies not in the measure but in the persistence of the measure after the conditions that justified it have changed. Where an institution grows and the composition of the cases it inspects shifts with that growth — the number of audited suppliers tripling while the proportion of problematic suppliers falls — the same metric improves automatically, though what has moved is the base itself rather than the detection capability sitting on top of it.
The institutional cost accumulates first in resource allocation. A control function that reports high accuracy is positioned in budget discussions as a well-performing unit; its case for additional investment weakens, headcount lost to attrition goes unreplaced, and the calibration of its thresholds goes unexamined for years at a stretch. Resource therefore flows not toward the place demonstrating measurable success but toward the place whose measurement method is structurally disposed to generate the appearance of it. The second site of accumulation is contractual: for as long as a supplier quality screen reports a high hit rate, inspection provisions in purchase agreements are relaxed, acceptance criteria are narrowed to a sampling basis, and the cost of the defect that passed through surfaces further along the supply chain as a recall provision or a rework line item.
The third and generally most expensive site of accumulation is valuation. When the quality of a control environment is interrogated at the diligence table during a sale or investment process, the first document produced is typically the hit-rate report; the second question from an experienced review team concerns how the events the system missed came to be identified. Where the answer is that they surfaced upon customer complaint, or after the incident had already materialised, the conclusion available to the reviewer is that the control system possesses no independent detection capability and is instead classifying realised events retrospectively. The transactional consequence of that finding is rarely a headline price reduction; it registers instead as an expansion of the representation and warranty package, an increase in the escrow proportion, or a separately drafted post-closing condition addressing unknown defects — which is to say the risk is left with the seller rather than priced into the consideration.
The first component of a structural intervention is the disaggregation of the single metric. In every rare-event decision domain, the number reported should not be one but at least three: the proportion of events that actually occurred which had been flagged in advance, the proportion of flagged cases that turned out to be genuine, and the base rate for the period itself. The third figure carries particular weight, since the interpretation of the first two depends entirely on knowing the base; where the base goes unreported, genuine improvement in detection and mere dilution of incidence become indistinguishable from one another. Presenting these three side by side, in the same table and at the same reporting frequency, produces a structurally different reading than presenting them in separate sections of the same pack.
The second component is pricing the cost of each error type before the decision rather than after it. The cost to an institution of a missed event and the cost of a false alarm are almost never equal; in occupational safety the ratio may differ by an order of magnitude, whereas in procurement screening it can invert entirely. Where threshold calibration follows the ratio between those two costs, the question of how much the system will be permitted to err ceases to be a technical parameter and becomes an explicit management decision with an owner attached to it. The third component is recording threshold changes in the decision log: where the sensitivity of an alerting system is reduced, and no written record captures who reduced it, on what reasoning and under what cost assumption, the metric improvement observed in the following period will be read as performance rather than as the arithmetic consequence of a deliberate narrowing.
Across the capital-intensive projects BEIREK manages, this intervention is made during the establishment of the project control architecture rather than after the first reporting cycle has set expectations. As early-warning indicators are defined, two distinct thresholds are set for each indicator — one triggering intervention, the other entering reporting — and the gap between them is calculated against what a missed event would cost in schedule terms and in cash flow terms. The same discipline governs contractor performance monitoring and procurement quality control: a clean result produced by any control point does not enter the project report until the case volume and the base rate from which that result emerged are stated alongside it.
Second, in every engagement the performance of the control system is itself tested at regular intervals through a retrospective miss review: adverse events that materialised during the period are taken one by one, and for each the record captures whether the system had any realistic opportunity to flag the event in advance and, where it did not, which data gap or threshold setting foreclosed that opportunity. This review functions as a calibration mechanism rather than an attribution exercise; its output is a revision to indicators and thresholds, not an assessment of individuals. Fixing the rhythm of that review on a monthly or quarterly cadence balances the control function's structural incentive to optimise its own headline metric against a standing obligation to report the boundaries of its own blind spot.
The rise of a measure and the rise of a capability are two separate phenomena, and in rare-event domains they move in opposite directions with some regularity. The maturity of an institution's control environment is evidenced not by the height of the hit rate it reports but by whether its own reporting discloses the base from which that rate is drawn and states plainly what the rate does not measure. For any performance figure arriving at the board table, the first question worth asking is not how high the number is, but what the same number would read on a system that did nothing at all.
