---
title: "Which Definition of Fairness: The Unrecorded Trade Against Accuracy in Scoring Systems"
description: "The fairness–accuracy trade cannot be eliminated, because equal pass rates, calibrated scores, and balanced error rates cannot hold simultaneously once group base rates differ. What remains governable is where the trade is made and how much it costs: the fairness definition is declared before scoring, the forfeited accuracy is measured, and every exception is recorded alongside the authority level that approved it."
url: https://www.beirek.com/en/blog/fairness-accuracy-trade-off
canonical: https://www.beirek.com/en/blog/fairness-accuracy-trade-off
published: 2025-04-26
modified: 2025-04-26
category: "Organisational Psychology"
category_url: https://www.beirek.com/en/blog/category/organisational-psychology
language: en-US
reading_time_minutes: 8
publisher: BEIREK LLC
publisher_url: https://www.beirek.com
license: "© BEIREK LLC — citation with attribution and link permitted"
keywords: ["fairness accuracy trade-off","scoring matrix governance","contractor prequalification","decision record discipline","due diligence documentation"]
topics: ["Organisational decision architecture","Vendor and contractor selection systems","Algorithmic fairness in institutional scoring","Diligence exposure from undocumented decisions"]
alternate_language_url: https://www.beirek.com/tr/blog/fairness-accuracy-trade-off
---

# Which Definition of Fairness: The Unrecorded Trade Against Accuracy in Scoring Systems

> **In short:** The fairness–accuracy trade cannot be eliminated, because equal pass rates, calibrated scores, and balanced error rates cannot hold simultaneously once group base rates differ. What remains governable is where the trade is made and how much it costs: the fairness definition is declared before scoring, the forfeited accuracy is measured, and every exception is recorded alongside the authority level that approved it.

*Every scoring matrix an organisation operates carries an undeclared definition of fairness, and that definition quietly purchases some portion of total predictive accuracy. The trade itself is unavoidable; the institutional cost arises when it is made after the results are visible and left off the record.*

---

In a contractor prequalification committee, disaggregating the output of the scoring matrix by region, firm size, or prior working relationship tends to reveal materially different pass rates across the resulting segments, and the meeting in which that divergence is first noticed is, more often than not, the same meeting in which the weight attached to one of the criteria is reduced. The criterion that gets reduced is rarely the most contested one; it is the one hardest to defend in the room — the count of completed works at comparable scale, the length of uninterrupted banking reference, or the record of liquidated damages assessed under prior contracts — that is, a line item carrying the distribution of the past directly into the present. The minutes record the rationale as a finding that the criterion's discriminating power was limited, when the criterion produced discomfort precisely because its discriminating power was high. How much predictive strength was surrendered in exchange for the correction appears in no document at all.

A second pattern, discussed far less openly, concerns the asymmetry with which such corrections are applied. The same organisation that softens the matrix in selection processes reported to its lenders or exposed to regulatory inspection will frequently continue running the original matrix, unmodified, across surfaces that are never reported outward — internal resource allocation, team assignment, subpackage award. Two different definitions of fairness are thus operated on two different surfaces within the same period, and neither is committed to writing. Over time this duality institutionalises itself not as an inconsistency but as two separate customs; when the person who established both departs, what remains is a set of numbers stripped of any record of the definition under which they were produced.

The mechanism underlying this behaviour is the fairness–accuracy trade-off — the conflict between a chosen definition of fairness and total predictive accuracy — and its most consistently misread feature is that it is not a question of will or intent. When three properties are demanded of a scoring system at once — equal pass rates across groups, scores that carry the same meaning relative to realised outcomes, and error rates distributed evenly across groups — those three conditions cannot hold simultaneously so long as the observed base rates of the groups differ. Tightening one necessarily loosens another, and this follows from the arithmetic of measurement rather than from any deficiency in the model. The instruction to build a matrix that is both fair and accurate is therefore not an executable instruction but an unsolvable equation until the operative definition of fairness has been named.

It would nonetheless be an error to assume the trade is always genuine, and this distinction constitutes the functional face of the mechanism. The historical performance label on which scoring systems are trained is, in a great many settings, not a clean measure of capability but a trace of past allocation decisions: a subcontractor awarded the easier packages carries a cleaner delivery record, a manager assigned to more visible projects accumulates stronger appraisal scores. Accuracy measured against such a label is, in substance, a measure of how faithfully the prior allocation logic has been reproduced. Under these conditions a fairness constraint may lower measured accuracy while raising the hit rate against realised outcomes; the trade appears on the page without existing in fact. The problem lies in the absence of any test capable of distinguishing which of the two cases obtains.

For this reason, most arguments about accuracy turn out, on inspection, to be arguments about the target variable. In a contractor selection matrix, success may be defined as delivery on schedule, as low change-order cost, as a scarcity of post-completion warranty calls, or as brevity of the negotiation cycle — and those four definitions rank the same firm in four different positions. The choice of target variable is made in the guise of a technical decision when it is in fact a strategic declaration of what the institution treats as valuable, and it is typically made, without deliberation, at the level of the analyst assembling the matrix. An accuracy rate computed before the target has been fixed remains, however high it reads, an answer to an unidentified question.

In capital-intensive projects, the institutional counterpart of this ambiguity surfaces first in unexpected line items. A prequalification threshold relaxed after the results were visible accumulates not in the tender file but in second-year rework cost, in liquidated damages claims, in the additional premium demanded at insurance renewal, and in critical-path items that begin to slide within the programme. The same relaxation takes a further form on the contract administration side: an agreement signed with a counterparty admitted through a lowered threshold typically requires a broader security package, a tighter interim payment regime, and more intensive site supervision, and the cost of that incremental burden is booked against the operating budget rather than against the scoring decision. The link between the decision and its price is thereby severed at the level of the ledger.

At the diligence table the matter assumes a sharper form. What a buyer or lender requesting documentation of selection and evaluation processes typically receives is not the matrix itself but the matrix as applied, and the gap between what was applied and what was declared consists of a series of manual interventions whose reasoning was never recorded. Identification of that gap rarely converts into a direct valuation discount; more characteristically it produces a widening of the representations and warranties, the opening of a distinct indemnity heading covering hiring and vendor selection, an increase in the escrow ratio, and the addition of a remediation undertaking to the conditions precedent. The organisation ends up paying, in the closing negotiation and with interest attached, for a decision it declined to document.

The heavier cost over a longer horizon is the precedent effect. A single exception granted without written rationale reappears in the following cycle under the formula of how it was handled last time, and within two or three cycles it has effectively voided the declared logic of the matrix; at that point the organisation holds a procedure no one applies alongside a practice no one has written down. The system continues to look operational for as long as a founder or long-tenured executive carries the bridge between those two artefacts in personal memory, and that dependency is precisely the kind investors price. What determines a company's valuation is less the correctness of its decisions than its ability to demonstrate that the same decision could be produced again under the same rules.

The intervention that neutralises this tendency is decision architecture rather than individual awareness, and it separates into four components. The first is declaration of the fairness definition before scoring begins: which equality is to be preserved — pass rate, consistency of score meaning, or error distribution — is written in a single sentence before the matrix is built, and that sentence cannot be amended once results are visible, only reopened in a subsequent cycle. The second is a target variable protocol: the observation by which success is measured, the measurement window, and the party producing the measurement are held in a document separate from the matrix itself. The third is a trade ledger, in which the predictive power surrendered by each applied constraint is measured and recorded numerically, so that the trade ceases to be a topic of debate and becomes a line item. The fourth is an exception register, in which every decision departing from the matrix is entered at the moment it is taken, together with its rationale and the authority level that approved it.

A cadence sits above these four components, because base rates are not stationary. As portfolio composition, geography, contract type, or market conditions shift, a calibration valid in one period drifts systematically in the next, and that drift is better tied to the project's own milestones — tender package release, first draw, scope change threshold, handover to operations — than to an arbitrary annual review. Embedded within the same cadence is a counter-argument role: a person other than the author of the matrix is required, in each cycle, to select one rejected candidate and defend the rationale for that rejection in writing. The role exists not to alter outcomes but to keep the reasoning articulable.

BEIREK's intervention in this area begins by bringing the full set of score-based processes — contractor prequalification, subpackage award, local content commitments, team allocation — under a single decision-record discipline. Before the scoring matrix is constructed, the target variable and the fairness definition are fixed on two separate pages, every weight assigned in the matrix is justified by reference to those two pages, and any weight adjustment made after results are opened is recorded under a new version number with the forfeited predictive power quantified. Exceptions are held in a separate register, and the approval level for each exception is defined in proportion to the contract value of the transaction; a deviation a project director may grant at subpackage level is escalated to investment committee level in principal contractor selection.

The review cadence is anchored to the project's own thresholds rather than to the calendar — tender package close, first draw, scope change, handover to operations — and at each threshold the distribution produced by the matrix is set alongside realised field performance so that the magnitude of calibration drift is expressed numerically. The output of the exercise is not a compliance report but a decision file that can be handed, as it stands, to the counterparty's diligence team during closing; which definition of fairness the organisation selected, what that selection cost on the accuracy side, and at which authority level the decision was taken are all legible from the file itself. The hardest question about an institution's selection system is not whether it is fair, but whether it can demonstrate in writing which definition of fairness it chose and what it knowingly gave up in return.

## Key Points

- Once observed base rates differ across groups, group parity and calibration cannot be satisfied at the same time; this is an arithmetic constraint rather than a matter of preference or model quality.
- When the weights in a scoring matrix are adjusted after the results are seen, the organisation loses any measure of the predictive power it has surrendered and, with it, the ground on which the decision could be defended.
- Most disputes framed as accuracy disputes are, on closer inspection, disputes about the target variable: until what counts as success is fixed, an accuracy figure answers a question no one has stated.
- A single exception recorded without its rationale becomes precedent in the following cycle and, after two or three repetitions, quietly voids the declared logic of the matrix.
- In diligence, an undocumented selection model is typically priced not as a headline valuation discount but as broader representations and warranties, a dedicated indemnity heading, and a higher escrow ratio.

## Questions

### Can fairness and accuracy be optimised simultaneously in a scoring system?

Not while group base rates differ. Equal pass rates, scores that carry consistent meaning relative to realised outcomes, and evenly distributed error rates cannot all hold at once; tightening one necessarily loosens another. The constraint is arithmetic and independent of model quality. What is governable is not the elimination of the trade but the prior declaration of which definition has been chosen, together with measurement and recording of the accuracy surrendered.

### Which fairness definition should be selected in a vendor or hiring scoring model?

The choice follows from which error is more expensive. Where the cost of an incorrectly admitted counterparty materialises as rework, liquidated damages, and security calls, consistency of score meaning takes precedence; where a narrowing candidate pool creates strategic exposure, balance of pass rates comes forward. What proves decisive is not the definition itself but the fact that it is fixed in writing before scoring and left unamended once results are visible.

### Why does recording manual exceptions to a scoring model matter?

An exception granted without written rationale becomes precedent in the next cycle and, after a few repetitions, voids the declared logic of the matrix in practice. The organisation is then left holding a procedure no one applies and a practice no one has documented, with the bridge between them carried in a single person's memory. Recording the exception at the moment of decision, with its rationale and approval authority, is the lowest-cost mechanism for removing that dependency.

### How is undocumented selection and scoring priced during due diligence?

Generally not as a headline valuation discount but as a tightening of the transaction structure. Once an unexplained gap is identified between the declared matrix and the decisions actually taken, the typical consequences are broader representations and warranties, a dedicated indemnity heading covering selection processes, a higher escrow ratio, and a remediation undertaking added to the conditions precedent. The cost is not removed; it is deferred to the closing negotiation.

---

Source: https://www.beirek.com/en/blog/fairness-accuracy-trade-off
Publisher: BEIREK LLC — https://www.beirek.com
