---
title: "The Human in the Approval Box: The Gap Between Oversight Existing and Oversight Working"
description: "Adding human approval to an automated process does not by itself eliminate risk. As the system's hit rate rises, the reviewer shifts from independent verification to confirmation of the output. Oversight becomes a genuine control layer only where the approver holds an independent information source, sufficient time to object, and an authority structure in which rejection carries no career cost."
url: https://www.beirek.com/en/blog/human-in-the-loop-complacency
canonical: https://www.beirek.com/en/blog/human-in-the-loop-complacency
published: 2025-04-19
modified: 2025-04-19
category: "Organisational Psychology"
category_url: https://www.beirek.com/en/blog/category/organisational-psychology
language: en-US
reading_time_minutes: 8
publisher: BEIREK LLC
publisher_url: https://www.beirek.com
license: "© BEIREK LLC — citation with attribution and link permitted"
keywords: ["human-in-the-loop complacency","automation oversight","internal control effectiveness","residual risk assessment","due diligence control testing","approval workflow governance"]
topics: ["Organisational decision architecture","Automation bias and control design","Risk register calibration","Transaction diligence on internal controls"]
alternate_language_url: https://www.beirek.com/tr/blog/human-in-the-loop-complacency
---

# The Human in the Approval Box: The Gap Between Oversight Existing and Oversight Working

> **In short:** Adding human approval to an automated process does not by itself eliminate risk. As the system's hit rate rises, the reviewer shifts from independent verification to confirmation of the output. Oversight becomes a genuine control layer only where the approver holds an independent information source, sufficient time to object, and an authority structure in which rejection carries no career cost.

*When a human sign-off is attached to an automated process, the residual risk is assumed to be closed; yet the reviewer, observing the system perform correctly over time, quietly reallocates attention elsewhere. The distance between oversight existing and oversight functioning surfaces institutionally on contractual, insurance and valuation surfaces.*

---

In a file arriving at a credit committee, an analyst's signature sits beneath the output of an automated scoring engine, attesting that the result has been reviewed and found appropriate. Across every file that same analyst has signed over the preceding quarter, the system's recommendation and the final decision have coincided without exception. That coincidence is consistent with two entirely different states of the world: either the engine is genuinely accurate and the analyst is confirming it against independent judgement, or the analyst has learned that the engine is accurate and has, in practice, stopped verifying anything at all. The file itself carries nothing capable of distinguishing between the two, since the signature looks identical in either case. On the institutional risk register, however, both states occupy the same cell, in the same colour, carrying the same assigned residual risk.

The same pattern recurs in settings far removed from automation. On a production line, the operator stationed at the visual inspection point is tasked with scanning each batch as it passes; as the line's defect rate falls, the probability that the operator actually detects a defect falls with it, because sustained search behaviour that goes unrewarded is not something human attention is built to maintain over long intervals. In a financial reporting chain, the distance between the preparer of a schedule and its approver contracts into a purely formal signature distance once the schedule has closed cleanly for several consecutive periods. In a supplier prequalification process, what is being examined on a given vendor's sixth consecutive pass is no longer the vendor, but the outcome of the preceding examination.

The behaviour has a settled name in the automation and decision-support vocabulary — **human-in-the-loop complacency**, the treatment of a human's presence in the loop as equivalent to that human exercising oversight. Its mechanism operates on two layers. The first is an attention economy: the reviewer derives a reliability expectation from the observed track record of the system under review, and allocates cognitive resource elsewhere in proportion to how high that expectation runs. This allocation is rational on its own terms, since spending scarce attention on a source that has not generated errors is, in the ordinary case, waste. The second layer is anchoring: once the system produces a number, a score or a recommendation, the reviewer's own estimate is no longer independent, forming instead within a narrow band around the displayed output. The human does not verify the figure; the human constructs a rationale consistent with it.

It matters to recognise that this is a shortcut rather than an error. An oversight layer that reconstructs every output from first principles roughly doubles the cost of the system it supervises and is, under most conditions, economically indefensible. So long as the system remains reliable, the reviewer's redeployment of attention raises aggregate organisational productivity. The difficulty lies not in the shortcut but in its persistence after the conditions that justified it have moved: when the input distribution drifts, when a model is retrained, when a supplier changes hands, when a regulatory threshold is revised, the system's hit rate deteriorates while the reviewer's confidence in it remains where it was. Confidence tracks performance with a lag, and within precisely that lag interval, the protection the oversight layer is presumed to supply is not present.

The institutional consequence appears first on the contractual surface. Processes incorporating human approval construct an architecture that migrates liability from provider to purchaser; software vendors' limitation clauses typically place responsibility for the ultimate use of an output with the customer, resting on the premise that a qualified user has examined it. Where that premise fails to hold, the economic location of the risk does not change — liability still sits with the purchaser — but the assertion that the risk is being actively managed has no foundation beneath it. A comparable asymmetry emerges on the insurance side, since professional indemnity pricing is built on the declared control architecture, and in an incident where a declared control was not in fact operating, the indemnity negotiation shifts from a debate about policy scope to a debate about the accuracy of the declaration itself.

A second consequence accumulates in internal control documentation. In audit trails, an approval step is measured by completion rate, and completion approaching one hundred per cent is reported as evidence that the control is healthy. The single statistic that speaks to that control's effectiveness, however, is not completion but the **rejection rate**. An approval step that has declined no output over twelve months either supervises a flawless process or is not a control at all but a recording action; attributing residual risk reduction to it without separating those two possibilities renders the risk matrix more favourable than the underlying reality supports. This is less a finding that internal audit uncovers than a blind spot generated by internal audit's own methodology.

A third consequence surfaces in sale processes and in valuation. Where a company presents automation intensity as a source of value, the buy-side diligence team asks not whether the system is accurate but what happens in the moment it is wrong; and when that question cannot be answered at document level, the valuation conversation migrates rapidly into representation and warranty scope, conditions precedent and escrow proportion. An oversight function described by a founder or a key technical employee in the form of "if the system produces a wrong output, I would notice" is, by definition, an oversight function that cannot be transferred, and a buyer prices it not as a control but as a key-person dependency. The institutional value of a process rests on demonstrating that it is repeatable independently of the founder, and the quality of human approval is among the weakest links in that demonstration.

The structural remedy does not run through asking reviewers to be more attentive, since attention is not a resource that responds to being requested. Effective intervention separates into three components. The first is **informational independence**: the approver is supplied with at least one data point not derived from the output under review — the raw input item alongside the model score, the site record alongside the supplier report, the source reconciliation alongside the preparer's schedule. The second is **time allocation**: the duration of the approval step is measured separately from the remainder of the process, and a median approval time falling below a defined threshold constitutes an early, quantifiable signal that the control has become formal. The third is **removing the cost of refusal**: where an approver's rejection delays the schedule and that delay is written into a performance indicator attributed to them, refusal behaviour is systematically suppressed, and the delegation architecture must therefore record a rejection as an output rather than a defect.

A fourth component is discussed less often but is the most determinative: **calibration testing**. Whether the oversight layer functions can only be established by injecting known-defective outputs into the process and measuring the rate at which they are intercepted. The purpose is not to examine the reviewer but to quantify the residual risk reduction actually delivered by the oversight layer, with the result forming the empirical basis for the control effectiveness score recorded in the risk matrix. In an organisation that runs no such test, the risk reduction attributed to the human approval step remains an assumption rather than a measurement.

BEIREK fixes this distinction at document level across the decision chains it manages in capital-intensive projects. For every automated or model-based output, three fields are maintained as a discrete element of the process map: the independent input on which the approval step rests, the authority level at which the approver holds a right of refusal, and the party to whom the schedule impact of a refusal is attributed. An approval record accordingly carries not only "who, and when" but "on the basis of what". On the same principle, the rejection rate and the median time expended on approval are tracked periodically for each approval step; where those two indicators decline together, the associated control is treated as having become formal, and its effectiveness is re-graded in the risk register.

On the diligence and investment committee side, the corresponding discipline is to test the control architecture presented rather than to accept it as described. For each human approval step appearing on a counterparty's control listing, the questions asked are how many refusals were issued over the preceding twelve months, on what stated grounds, and how the process operated following a refusal; steps for which no answer is available are modelled not as controls delivering risk reduction but as recording actions, with residual risk repriced accordingly. That reclassification bears directly on the scope of conditions precedent, on the escrow proportion, and on the first ninety days of the integration plan, because whether the acquired item is a control or merely a signature field becomes visible only once the question is put.

Oversight lives in a behaviour rather than in an organisational chart, and the existence of that behaviour is evidenced not by the approval box being filled but by a demonstration that the box has, at least once, been left empty. An approving authority that has never said no conveys no information when it says yes.

## Key Points

- As a system's accuracy improves, the human approver's capacity for independent verification tends to decline, meaning reliability and oversight effectiveness move in opposite directions rather than together.
- The existence of an approval record is not evidence that a control is operating; the trace of a functioning control lies not in the count of approvals but in a rejection and correction rate that is measurably different from zero.
- Human sign-off is institutionally useful to the extent that it relocates liability from the system to an accountable individual, but relocation is not reduction — the contractual position of the risk changes while its probability does not.
- Oversight works only where the approver is supplied with a second information source not derived from the output under review, is given enough time to form a judgement, and operates under a delegation structure in which refusal is recorded as an output rather than a defect.
- In diligence, the discriminating question about human-approved processes is not whether the control exists but how many times, and on what stated grounds, it rejected an output over the preceding twelve months.

## Questions

### Does adding human approval genuinely reduce the risk of an automated system?

Only under specific conditions. An approval layer reduces risk where the approver holds an information source independent of the output under review, sufficient time to evaluate it, and genuine authority to refuse. Absent those three, the approval step is a recording action rather than a control; it relocates liability from the provider to the institution without altering the probability of error.

### How can it be established whether an approval step has become merely formal?

Two indicators are read together: the rejection rate and the median time expended on approval. A step in which no output has been declined over an extended period, and in which time spent on approval is steadily falling, indicates not that the supervised process is flawless but that oversight has in practice been abandoned. Completion rate does not draw this distinction, since full completion appears identical in either state.

### How should human-approved controls be tested during due diligence?

The question is not whether the control exists but how many times it declined an output over the preceding twelve months. Where the number of refusals, their stated grounds, and the process that followed cannot be evidenced in documentation, that step is modelled as a recording action rather than a control delivering risk reduction. The reclassification bears directly on residual risk, escrow proportion and the scope of conditions precedent.

### Does training reviewers to be more attentive help?

Attention is not a resource produced on request; sustained search behaviour directed at a process that generates no errors is maintained through process redesign rather than instruction. Effective intervention consists of informational independence, separately measured approval time, removal of the cost of refusal, and periodic calibration testing using known-defective outputs. Awareness training does not substitute for any of those four.

---

Source: https://www.beirek.com/en/blog/human-in-the-loop-complacency
Publisher: BEIREK LLC — https://www.beirek.com
