---
title: "What a High Hit Rate Conceals: The Accuracy Paradox in Corporate Decision Metrics"
description: "In decision domains where the target event is rare, overall accuracy is misleading: a system that flags nothing at all still reports accuracy equal to the base rate. Meaningful measurement requires asymmetric metrics in which the cost of a missed event and the cost of a false alarm are priced separately, because a single percentage renders that asymmetry invisible."
url: https://www.beirek.com/en/blog/accuracy-paradox-decision-metrics
canonical: https://www.beirek.com/en/blog/accuracy-paradox-decision-metrics
published: 2025-04-25
modified: 2025-04-25
category: "Organisational Psychology"
category_url: https://www.beirek.com/en/blog/category/organisational-psychology
language: en-US
reading_time_minutes: 7
publisher: BEIREK LLC
publisher_url: https://www.beirek.com
license: "© BEIREK LLC — citation with attribution and link permitted"
keywords: ["accuracy paradox","base rate","class imbalance","early warning systems","control environment","false negatives","threshold calibration","due diligence"]
topics: ["Measurement architecture in corporate control functions","Rare-event detection and metric design","Control environment assessment in due diligence","Asymmetric error costs and threshold governance"]
alternate_language_url: https://www.beirek.com/tr/blog/accuracy-paradox-decision-metrics
---

# What a High Hit Rate Conceals: The Accuracy Paradox in Corporate Decision Metrics

> **In short:** In decision domains where the target event is rare, overall accuracy is misleading: a system that flags nothing at all still reports accuracy equal to the base rate. Meaningful measurement requires asymmetric metrics in which the cost of a missed event and the cost of a false alarm are priced separately, because a single percentage renders that asymmetry invisible.

*As the reported hit rate of a forecasting, alerting or screening system rises, the information it conveys to the institution may fall; in any decision domain dominated by rare events, an overall accuracy percentage measures the base rate rather than the capability, and the balance-sheet consequence of that distinction accumulates in the cost of the event that was never flagged.*

---

When a board presentation reports that the risk early-warning system operated at a ninety-seven percent hit rate over the preceding quarter, the reaction around the table is typically approval; the system is working, its budget is renewed, and the team's performance review closes favourably. In the same meeting, the fact that none of the three material incidents that actually occurred during that quarter had been flagged in advance sits under a separate agenda item, often in the annexes to the operations report. Both pieces of information reach the same room, at the same hour, in front of the same people, yet the logical relationship between them is never established, since the first has been classified as a performance metric and the second as an incident log. This separation is not a lapse of attention; to the extent that the reporting architecture itself files the two items in different sections, the work of connecting them is quietly delegated to the individual effort of whoever happens to be reading closely.

The same pattern recurs across credit risk scoring, supplier quality screening, internal audit sampling, fraud detection, occupational health and safety whistleblowing channels, and any classification model built on machine learning. The shared condition is straightforward: the event being measured is rare, while the state in which nothing happens is dominant. Under that condition, a system that flags nothing whatsoever — one that classifies every case as clean — will report ninety-seven percent accuracy where the underlying incidence of the event is three percent. The metric is arithmetically correct, the calculation withstands audit, and the number survives review without objection; what it carries, however, is not a statement about the capability of the system but a restatement of how uncommon the event happens to be.

The structure at work here is the accuracy paradox — the condition under which a rising overall hit rate can coincide with a falling capacity to find the very thing the system exists to find. The mechanism operates on two layers. The first is arithmetic: under an imbalanced distribution, selecting the majority class in every instance produces the highest score when performance is viewed through a single metric, so any system operating under optimisation pressure will drift naturally in that direction. The second layer is organisational and, in practice, more determinative: once a metric is attached to performance appraisal, budget allocation or a management report, raising that metric becomes rational behaviour for the team responsible for it, and in a rare-event domain the cheapest available route to raising it is to lift the threshold, which is to say to flag less.

Reading this tendency as a defect would misstate the case. An overall accuracy figure is a highly functional summariser under conditions in which the classes are balanced and the costs of the two error types sit close to one another; its reduction to a single number is precisely what permits comparison, time-series tracking and benchmarking across business units that otherwise share no common language. The difficulty lies not in the measure but in the persistence of the measure after the conditions that justified it have changed. Where an institution grows and the composition of the cases it inspects shifts with that growth — the number of audited suppliers tripling while the proportion of problematic suppliers falls — the same metric improves automatically, though what has moved is the base itself rather than the detection capability sitting on top of it.

The institutional cost accumulates first in resource allocation. A control function that reports high accuracy is positioned in budget discussions as a well-performing unit; its case for additional investment weakens, headcount lost to attrition goes unreplaced, and the calibration of its thresholds goes unexamined for years at a stretch. Resource therefore flows not toward the place demonstrating measurable success but toward the place whose measurement method is structurally disposed to generate the appearance of it. The second site of accumulation is contractual: for as long as a supplier quality screen reports a high hit rate, inspection provisions in purchase agreements are relaxed, acceptance criteria are narrowed to a sampling basis, and the cost of the defect that passed through surfaces further along the supply chain as a recall provision or a rework line item.

The third and generally most expensive site of accumulation is valuation. When the quality of a control environment is interrogated at the diligence table during a sale or investment process, the first document produced is typically the hit-rate report; the second question from an experienced review team concerns how the events the system missed came to be identified. Where the answer is that they surfaced upon customer complaint, or after the incident had already materialised, the conclusion available to the reviewer is that the control system possesses no independent detection capability and is instead classifying realised events retrospectively. The transactional consequence of that finding is rarely a headline price reduction; it registers instead as an expansion of the representation and warranty package, an increase in the escrow proportion, or a separately drafted post-closing condition addressing unknown defects — which is to say the risk is left with the seller rather than priced into the consideration.

The first component of a structural intervention is the disaggregation of the single metric. In every rare-event decision domain, the number reported should not be one but at least three: the proportion of events that actually occurred which had been flagged in advance, the proportion of flagged cases that turned out to be genuine, and the base rate for the period itself. The third figure carries particular weight, since the interpretation of the first two depends entirely on knowing the base; where the base goes unreported, genuine improvement in detection and mere dilution of incidence become indistinguishable from one another. Presenting these three side by side, in the same table and at the same reporting frequency, produces a structurally different reading than presenting them in separate sections of the same pack.

The second component is pricing the cost of each error type before the decision rather than after it. The cost to an institution of a missed event and the cost of a false alarm are almost never equal; in occupational safety the ratio may differ by an order of magnitude, whereas in procurement screening it can invert entirely. Where threshold calibration follows the ratio between those two costs, the question of how much the system will be permitted to err ceases to be a technical parameter and becomes an explicit management decision with an owner attached to it. The third component is recording threshold changes in the decision log: where the sensitivity of an alerting system is reduced, and no written record captures who reduced it, on what reasoning and under what cost assumption, the metric improvement observed in the following period will be read as performance rather than as the arithmetic consequence of a deliberate narrowing.

Across the capital-intensive projects BEIREK manages, this intervention is made during the establishment of the project control architecture rather than after the first reporting cycle has set expectations. As early-warning indicators are defined, two distinct thresholds are set for each indicator — one triggering intervention, the other entering reporting — and the gap between them is calculated against what a missed event would cost in schedule terms and in cash flow terms. The same discipline governs contractor performance monitoring and procurement quality control: a clean result produced by any control point does not enter the project report until the case volume and the base rate from which that result emerged are stated alongside it.

Second, in every engagement the performance of the control system is itself tested at regular intervals through a retrospective miss review: adverse events that materialised during the period are taken one by one, and for each the record captures whether the system had any realistic opportunity to flag the event in advance and, where it did not, which data gap or threshold setting foreclosed that opportunity. This review functions as a calibration mechanism rather than an attribution exercise; its output is a revision to indicators and thresholds, not an assessment of individuals. Fixing the rhythm of that review on a monthly or quarterly cadence balances the control function's structural incentive to optimise its own headline metric against a standing obligation to report the boundaries of its own blind spot.

The rise of a measure and the rise of a capability are two separate phenomena, and in rare-event domains they move in opposite directions with some regularity. The maturity of an institution's control environment is evidenced not by the height of the hit rate it reports but by whether its own reporting discloses the base from which that rate is drawn and states plainly what the rate does not measure. For any performance figure arriving at the board table, the first question worth asking is not how high the number is, but what the same number would read on a system that did nothing at all.

## Key Points

- In rare-event decision domains, an overall hit rate reports the base rate rather than the performance of the system, and the two remain uninterpretable until they are reported separately.
- The cost of a missed event and the cost of a false alarm are almost never symmetric, yet a single accuracy percentage pushes that asymmetry outside the measurement frame entirely.
- An alerting system that shifts its threshold toward flagging less will report progressively higher accuracy, which reflects an expanding blind spot rather than an improving capability.
- The structural remedy is measurement architecture rather than individual scepticism: separate reporting of the base rate, advance pricing of the cost of a miss, and a written decision record for every threshold change.
- At the diligence table, a control system is assessed not by the number of events it caught but by how the events it missed were eventually detected.

## Questions

### What is the accuracy paradox, and why does a high accuracy figure conceal a weak system?

Where the target event is rare, a system that classifies every case as clean still produces accuracy equal to the base rate. If the incidence of the event is three percent, a model that flags nothing reports ninety-seven percent accuracy. The metric is mathematically sound and will survive audit, but the information it carries concerns the rarity of the event rather than the capability of the system, which is why it cannot stand alone as a performance indicator.

### How can it be established whether an early-warning system is genuinely working?

Three figures should be reported together rather than one percentage: the proportion of realised events flagged in advance, the proportion of flagged cases that proved genuine, and the base rate for the period. Without the base rate the first two cannot be interpreted, since improvement in detection is indistinguishable from dilution of incidence. Presenting the three in a single table at a single frequency yields a materially different reading than distributing them across separate sections.

### How is the quality of a control system assessed during due diligence?

The review proceeds not from the hit-rate report submitted but from the question of how the events the system missed were eventually identified. Where the answer is upon customer complaint, or after the incident materialised, the system has no independent detection capability and is classifying realised events retrospectively. The transactional consequence is typically not a price reduction but an expanded representation and warranty package or an increased escrow proportion.

### How should the balance between false alarms and missed events be set?

That balance is an explicit management decision rather than a technical parameter. The cost of each error type is priced separately in advance — in occupational safety the cost of a miss may be an order of magnitude heavier, whereas in procurement screening the ratio can invert — and the threshold is calibrated against that ratio. Every subsequent threshold change should enter the decision record together with its author, its reasoning and its underlying cost assumption.

---

Source: https://www.beirek.com/en/blog/accuracy-paradox-decision-metrics
Publisher: BEIREK LLC — https://www.beirek.com
