---
title: "What Rare Events Never Teach: Class-Imbalance Bias in Institutional Decision Systems"
description: "Class-imbalance bias is the systematic under-learning of infrequent outcomes because they occupy only a small share of the observation pool, a failure shared by statistical models and by managerial judgement alike. Its institutional consequence is a decision system that is highly accurate on average and blind in the tail; neutralising it requires treating the rare class under a separate evidentiary regime."
url: https://www.beirek.com/en/blog/class-imbalance-bias-rare-event-decisions
canonical: https://www.beirek.com/en/blog/class-imbalance-bias-rare-event-decisions
published: 2025-04-25
modified: 2025-04-25
category: "Organisational Psychology"
category_url: https://www.beirek.com/en/blog/category/organisational-psychology
language: en-US
reading_time_minutes: 8
publisher: BEIREK LLC
publisher_url: https://www.beirek.com
license: "© BEIREK LLC — citation with attribution and link permitted"
keywords: ["class-imbalance bias","tail risk in capital projects","decision architecture and approval thresholds","false negative cost asymmetry","survivorship in institutional records","covenant and reserve calibration"]
topics: ["Organisational decision-making under skewed outcome distributions","Risk calibration in project finance and contingency sizing","Governance design for counter-argument and escalation"]
alternate_language_url: https://www.beirek.com/tr/blog/class-imbalance-bias-rare-event-decisions
---

# What Rare Events Never Teach: Class-Imbalance Bias in Institutional Decision Systems

> **In short:** Class-imbalance bias is the systematic under-learning of infrequent outcomes because they occupy only a small share of the observation pool, a failure shared by statistical models and by managerial judgement alike. Its institutional consequence is a decision system that is highly accurate on average and blind in the tail; neutralising it requires treating the rare class under a separate evidentiary regime.

*An organisation's historical record represents its most expensive outcomes with the fewest observations, and that asymmetry causes both statistical models and managerial intuition to systematically under-learn the rare case. The result is a decision architecture that performs with high accuracy in ordinary conditions and falls silent precisely when the outcome matters most.*

---

In an investment committee session reviewing five years of performance across a development portfolio, the substance of the presentation is almost invariably built from completed assets — realised returns, variance against budget, schedule performance measured against the approved baseline — while projects that were cancelled, that failed during permitting, or that never reached financial close appear as a single line, expressed as a count rather than as a distribution. When the same committee turns, minutes later, to the next allocation decision, the reference set alive in the room consists entirely of those completed assets; the rare but expensive outcomes are present neither as a distribution nor as a recognisable pattern. The same asymmetry is observable in credit committees, in supplier prequalification, and in hiring panels: the decision maker knows the mechanics of the frequent outcome in fine detail and knows the infrequent outcome only by name. This arises from the structure of the sample rather than from inattention, which is exactly why attentiveness alone cannot correct it.

The identical asymmetry surfaces in more measurable form wherever the decision has been delegated to a system. A supplier-default prediction model trained on thousands of purchase orders containing a few dozen defaults can produce an overall accuracy above ninety-nine per cent by never predicting a default at all, and where aggregate accuracy is the only figure reported upward, that model reads as a success. It is not a success; it has memorised the majority class and has never once performed the task it was commissioned to perform, which was to flag the exceptional case. Human judgement and statistical estimation converge here on the same failure for the same reason, and that convergence is itself the diagnostic point: the problem does not live in the software layer, and replacing the model with a better-specified one leaves it untouched. What is at issue is the evidentiary regime under which the rare outcome is permitted to enter the record.

The pattern carries a name — class-imbalance bias, the condition in which one outcome category occupies so overwhelming a share of the observation pool that the minority category becomes neither learnable nor testable. The mechanism operates in two layers. In the first, any learning process, whether it is the parameter calibration of an algorithm or the intuition a manager accumulates across a career, orients itself toward minimising total error, and ignoring the rare class outright is the cheapest available route to a smaller error total. In the second, a feedback loop closes: because the rare outcome is not anticipated, no countermeasure is designed against it; because no countermeasure existed, its eventual occurrence is coded after the fact as a one-off misfortune attributable to circumstances that will not recur; and because it has been coded as singular, it never enters the forward sample as a pattern that could inform the next decision.

There is a range of conditions under which this tendency is genuinely efficient, and designing an intervention without acknowledging that range begins from the wrong premise. Where the outcome distribution is stable, where the cost of the rare event remains within the same order of magnitude as the cost of the ordinary event, and where decisions recur at high frequency, a rule calibrated to the majority class is a sound economising device; examining every low-probability branch individually burns more in review cost than the examination returns. The difficulty lies not in the shortcut but in its persistence after the conditions that justified it have lapsed. In a capital-intensive project where the tail outcome carries a cost an order of magnitude greater than the typical outcome, a frequency-weighted decision rule is no longer efficient — it is simply uncalibrated, and its apparent economy is borrowed against a loss that has not yet been recognised.

The institutional cost appears first in budget and schedule assumptions. Where the duration estimate for a construction programme is derived from the mean duration of completed works, the events that constitute the tail — a permitting objection escalated into litigation, a slipped delivery date on a single-source long-lead item, an interconnection study returned for re-run — never enter the estimating distribution at all, and the contingency line intended to absorb programme risk is sized against the observed body of that distribution rather than against a tail that has never been characterised. This sizing error does not present itself as a discrete entry anywhere in the accounts; it emerges in fragments, through elevated working capital requirements, through unscheduled draw requests, and through exposure to liquidated damages that was priced as remote. A cash flow model can misstate nothing at the level of individual line items and still rest on a wholly incorrect distributional assumption.

The second cost accrues in the contractual and financing layer. Where DSCR tests, reserve account sizing, and covenant thresholds are calibrated against historical variance, and that historical variance excludes the rare class by construction, the resulting thresholds run systematically loose; the cushion visible to the lender is a cushion that has never been tested against the scenario class most likely to consume it. The same logic governs representation and warranty negotiation: where the data room documents only disputes that materialised, leaving exposures that were structurally available but never realised without any documentary trace, warranty scope is negotiated down toward the observed event set, and a tail event surfacing after closing falls within no party's covenant. Escrow proportions indexed to a claims history compiled on the same basis inherit precisely the same blind spot, and inherit it without any party having made an identifiable error.

The third cost sits inside institutional learning itself and is the least visible of the three. Where an organisation retains records only for the projects that survived, its archive becomes progressively less informative over time, because every additional successful file deposited into it increases the weight of the majority class and further depresses the relative visibility of the rare one. A firm with a decade of institutional memory, if that memory consists exclusively of completed work, is not better informed about tail risk than a firm with three years of history; it is merely more confident. At the diligence table this distinction can be read against institutional maturity rather than in its favour, since what a reviewer is looking for is not the volume of accumulated experience but the demonstrated capacity to disaggregate that experience by risk class — a capacity that a survivor-only archive is structurally unable to evidence.

The first component of a structural intervention is to place the rare class under a separate evidentiary regime, so that tail events are not pooled into the same statistic as body events but held in a distinct record under a distinct classification. The second component is asymmetric threshold construction: where the cost of a false negative exceeds the cost of a false positive by an order of magnitude, the decision rule is shifted to reflect that ratio explicitly, and the reasoning behind the shift is committed to writing rather than left as a matter of judgement. The third is the systematic capture of outcomes that did not occur — bids submitted and lost, projects screened and declined, contracts negotiated and abandoned before signature — since these constitute a sample in their own right, and without that sample nothing about the rare class can be learned. The fourth is the relocation of the reporting metric away from aggregate accuracy and onto the capture rate over the rare class.

BEIREK embeds this intervention inside the project management line rather than delivering it as a separate analytical product. In the programmes we run, the decision record opens at the moment of proposal rather than at the moment of approval, so that the basis on which an assumption was selected — the sample it rested on, the alternative that was rejected, the reasoning that connected the two — is committed to writing before the outcome is known, which prevents the outcome from retroactively reshaping the account of how the decision was reached. Alongside this, files that were declined or that lapsed are maintained under the same documentary discipline as files that completed, on the reasoning that the true risk distribution of a portfolio becomes visible only when the unclosed files are counted as observations rather than treated as absences.

Our second line of intervention defines tail scenarios by condition rather than by probability. Instead of estimating how often a given event will occur, we specify which conditions, satisfied jointly, render that event available, and then bind each condition to an observable indicator — the expiry of an objection period in a permitting file, a second deferral of a confirmed delivery date from a single-source supplier, a scope revision request within an interconnection study. What follows when an indicator triggers is not an alert circulated for awareness but a predefined review track operating under a different approval threshold, with a named owner and a defined response window. Pre-mortem sessions form part of this line, and their agenda is not the project's route to success but the mechanism through which its failure would have occurred; the output is not a risk register but a specification of indicators and of the procedure that runs when one of them fires.

Managing class imbalance is, in the end, a question of authority rather than a question of data. The signal that points toward the rare class runs, by definition, against the accumulated experience of the majority, and the individual or function carrying that signal occupies a structurally weak position relative to the decision authority that embodies majority experience; the attenuation of such signals reflects the distribution of authority within the institution rather than any deficit of individual courage. For that reason, unless it has been settled in advance who plays the counter-argument role, under what mandate, and at which stage of the process, no measurement correction will suffice on its own. The question worth putting to an institution is therefore not how prepared it is for the rare event, but whether the person who raised the rare event is taken seriously in the room without first having to be proved right.

## Key Points

- A very high overall accuracy rate is readily achievable by a rule that never predicts the rare outcome at all, which is why aggregate accuracy is not, on its own, a measure of decision performance.
- The under-representation of rare events is a property of the events rather than a defect in the data, so the corrective is not more data but an asymmetric decision threshold that reflects the cost differential.
- Where an institution's archive contains only projects that survived, the sample available for reasoning about tail risk remains structurally incomplete no matter how long the archive runs.
- When the cost of a false negative exceeds the cost of a false positive by an order of magnitude, keeping the decision rule symmetric is not a deliberate choice but an uncalibrated assumption.
- Managing the rare class is institutionalised through a separate review track and a distinct approval threshold, not through individual vigilance on the part of the decision maker.

## Questions

### What exactly is class-imbalance bias?

It is the under-learning of a minority outcome category because one category occupies an overwhelming share of the observation pool. The mechanism operates identically in statistical models and in managerial intuition, since the cheapest available route to a smaller total error is to disregard the rare class altogether. The institutional result is a decision system that performs accurately under ordinary conditions and remains silent in the tail, where the consequences are largest.

### Our model reports very high accuracy — can this still be a problem?

Yes, and a very high accuracy figure is frequently the symptom rather than a defence against it. Where the rare event occupies a small share of the sample, a rule that never predicts that event will nonetheless post an excellent aggregate score. The meaningful measures are the capture rate over the rare class and the cost-weighted total of false negatives; aggregate accuracy reports mainly on how skewed the underlying sample happens to be.

### How is risk analysed when data on rare events is scarce?

By substituting condition specification for probability estimation. Rather than estimating how frequently an event will occur, the analysis records which conditions, satisfied together, make it available, and binds each condition to an observable indicator. When an indicator triggers, a predefined review track and a different approval threshold come into effect. The decision then rests on a monitorable structure rather than on a statistic that the sample was never capable of supporting.

### How does this bias affect company valuation?

Where historical variance excludes the rare class, DSCR tests, reserve sizing, and covenant thresholds are calibrated loose, and the cushion a lender observes has never been tested against the scenario class most likely to consume it. The same blind spot carries into warranty scope and escrow proportions. Diligence teams also look past the volume of accumulated experience toward whether that experience can be disaggregated by risk class at all.

---

Source: https://www.beirek.com/en/blog/class-imbalance-bias-rare-event-decisions
Publisher: BEIREK LLC — https://www.beirek.com
