The economics of exceptions
Why AI should follow damage, not cost.
The argument in brief
Industrialisation made repetition cheap. Digitalisation made coordination cheap. AI makes judgment cheap. Each wave collapsed the dominant cost of getting something done, and each time the residual moved somewhere else.
A standard case is repetition and coordination. That is why the standard case has been getting cheaper for a century. An exception is judgment, which is why it never did.
AI changes the economics of exceptions. It does not much change the economics of average work, because the previous two waves already did that.
Judgment has been purchasable before, through expert systems, decision support tools, specialist consultants. It was never cheap enough to spend on every case, so firms spent it only on cases that obviously mattered and suppressed the rest. Tighter policies. Narrower terms. Computer says no. The exception rate came down and the exceptions that survived stayed stubbornly human.
Cheap judgment removes the reason to suppress. That makes the residual the whole opportunity, and it relocates the investment question. The right test is not which process costs most to run. It is which process carries most value at risk. Firms know the first number to the decimal place and have never calculated the second.
The clerk we fired made the same argument one layer down, about the individual. Administrative cognition was pushed onto the desks of people paid for judgment. This paper is the institutional version: damage pushed onto the path nobody designed. Different subject, same displacement.
Why the damage sits in the exception
The standard route through a process is designed. Someone specified it, tested it, instrumented it, wrote controls around it, trained people to follow it. Failures there are anticipated failures, which is why they are cheap: the invoice gets reissued, the match gets cleared, the ticket closes by Thursday.
The exception route is the route nobody designed. It exists because a case arrived that the specification did not imagine. No control, no test, no instrumentation, often no record beyond an inbox.
A firm's entire tail risk lives on the path it never drew.
In regulated industries it is worse than undesigned. It is negotiated. Service level agreements specify performance on the standard case and then carve out the conditions under which the obligation lapses. UK telecoms has a name for the carve-out, MBORC, matters beyond our reasonable control, which suspends targets exactly when customers are least able to tolerate failure. The regulator and the firm jointly specified the happy path and jointly declined to specify the rest.
Where someone has specified the standard case, assume nobody specified the exception. That is falsifiable, and testable in any regulated industry in an afternoon.
The internet in the rain argued that copper service levels were never engineered to succeed. They were an economic settlement between state and company, dimensioned to disappoint, because dimensioning for every storm builds a network nobody can afford. That is the exception path as deliberate policy rather than oversight.
Why firms cannot see this
The blindness is structural, not stupid. It follows from what the management system measures.
Finance counts the cost of running a process every month: headcount, cost to serve, cost per transaction. Visible, recurring, comparable across periods, owned by somebody. A programme that reduces them produces a number a CFO can verify in ninety days.
The cost of the exception is measured nowhere, because most of it has not landed yet. There is no line item for the restatement avoided, the tribunal that never convened, the price increase the market would still have accepted. Absence has no dashboard.
Firms do not underinvest in tail risk because they misjudge it. They underinvest because their instruments do not point at it, and nobody was ever promoted for a crisis that did not occur.
Two ways a process goes wrong
Consider seven common end-to-end processes cutting across a firm's departments: Order to Cash, Trouble to Resolve, Concept to Market, Procure to Pay, Hire to Retire, Record to Report, Forecast to Plan. Each has a standard route and two ways of leaving it.
The first is too much work at once. A storm takes out a region and the contact centre takes ten times its normal call volume on a Tuesday. Nothing is wrong with the process. There is more of it than there are people.
The second is a case the process was never designed to handle. A customer disputes an invoice spanning two contracts, three tax jurisdictions and a credit hold that should have expired. No standard route exists, so it goes to a person, who spends four hours on it.
The first is a staffing problem. The second is a design problem. Only the second is where AI can shrink the stock of unresolved work rather than staff around it.
Against the first, AI forecasts better: more data, finer patterns, earlier warning of the Tuesday spike. Worth having, and capped. No model predicts a storm nobody saw or a regulator's surprise announcement. Better forecasting improves how a firm staffs against a fixed problem. It does not shrink the problem.
Most AI spending today addresses the first problem in a third way, neither forecasting the spike nor redesigning the process but absorbing the volume directly. The AI contact centre agent adds capacity that needs no hiring, and in many deployments substitutes for capacity the firm already has. That case is real, it is well understood, and the vendor market serves it competently. It is not the argument here. This paper takes no position on how much of a firm's standard-case capacity should be human, because that question is already being asked in every boardroom.
Capacity substitution changes what the standard case costs. Exception recovery changes what the firm can absorb. The first is a labour question. The second is a design question.
Against the second, AI spots the case coming, handles it with judgment and rules instead of a four-hour escalation, and remembers how. Each unusual case resolved becomes a case the system handles by itself next time. The set of things requiring a human contracts, permanently.
Forecasting gives a better estimate every year. Recovery gives a smaller problem every year. Over five years these are not comparable.
Not all exceptions are the same
The recovery claim needs a distinction the word exception hides, because without it the argument does not survive its own definition. If an exception is the case the specification did not imagine, then it is precedent-free, and precedent-free is exactly where machine learning is weakest. A system that learns from resolved cases cannot learn from a case that has never occurred.
The resolution is that most of what a firm calls an exception is not novel at all. It is recurring and unspecified, which is a different thing entirely.
Three categories, not two. The specified and recurring case, which every wave of automation has already made cheap. The recurring but unspecified case, which nobody ever wrote down. And genuine novelty, which has no precedent and will not have one next quarter either.
The middle category is the prize, and it is far larger than firms assume. The invoice spanning two contracts and three tax jurisdictions is not unprecedented. It happens forty times a year. Nobody specified it because forty cases never justified the cost of a specification, an integration and a trained team, and because judgment was expensive enough that handing it to a person was the rational answer each time. The category exists because of a cost threshold, not because of genuine unpredictability.
Cheap judgment moves the threshold. Work that was never worth specifying becomes work worth handling, and that is the whole of the opportunity.
Genuine novelty stays human, and should. A case with no precedent gives a model nothing to generalise from, and confident handling of it is the failure mode to fear rather than the capability to buy. The honest claim is narrower than it first appears and stronger for being bounded: AI does not make judgment cheap in general. It makes judgment cheap on the recurring case nobody specified, and that case is most of the volume and most of the damage.
The two numbers
Sequencing AI investment across the seven requires two numbers per process. Every firm has the first. No firm has the second.
Cost to run is already on the P&L. Headcount, cost to serve, cost per case, systems and licences allocated down. It takes a week to assemble because somebody assembles it every month anyway.
Value at risk is the cost of the exception path, and it has two components that behave differently and are usually confused.
Tail risk is one case producing one large event. Accreting friction is many cases producing no event at all, until the day the firm needs something from a counterparty it has spent years wearing down.
Tail risk is the recognisable form. A reconciliation break that becomes a restatement. A leaver dispute that becomes a tribunal. A product decision taken on a stale reading of the market that costs three years of capex. Rare, discrete, and large enough that the finance function would notice if anyone were counting.
Accreting friction is the form firms miss entirely. A customer meets the firm on the path with no script and no owner, and quietly revises what they will pay for the relationship. It shows up as nothing for years. Retention holds. Renewals tick over. Then it appears all at once, in the commercial line rather than the risk register: price tolerance gone, upsells dying in committee, renewals that used to close in a call now needing a discount.
Exception handling is not insurance against a disaster. It is the forward price of the revenue line.
This bites hardest at the stall. A firm that is compounding absorbs a great deal of accumulated resentment, because the numbers are carrying the argument. A firm that stops growing has to defend price, and price is defended with goodwill it either has or does not. The friction is invisible while it is being incurred and decisive at the moment it is needed.
At the far end of the same mechanism sits reputation, which is accreting friction that has outlived the fix. Enough badly-handled exceptions over enough years and a firm acquires a standing that survives the technology, the process and the people responsible. BT is one visible example among several. Nobody there today caused it. Everybody there today pays for it.
Customer emotional debt set out this mechanism: customers do not stay for value alone, the debt builds behind a dam while retention metrics hold, and it comes due as lost pricing power rather than as churn. This paper identifies where the debt is issued. It is issued on the exception path, because that is the only place the firm has no script.
How to get the second number
Value at risk cannot be calculated, and any figure produced to two decimal places is false precision. It can be bounded, process by process, in about six weeks. What comes out is an exposure range good enough to rank investment, not a number good enough to book.
For tail risk, pull the worst three failed cases in each process over the last five years and cost them fully: remediation, legal, regulatory, management time, and the revenue that moved as a result. Most firms have never assembled this because no single function owns it, and the assembling is itself the finding. Multiply by an honest recurrence view rather than a probability.
For accreting friction, compare the discount curve on accounts that have been through the exception path against those that have not, controlling for size and tenure. The data sits in the CRM and the billing system and has almost certainly never been joined. Read the gap carefully: difficult accounts both generate more exceptions and negotiate harder, so part of any gap is selection rather than damage. Split by whether the exception was the firm's fault to get closer to the causal share. A material gap that survives that split is the number worth acting on.
Expect the exercise to be resisted. It produces a figure nobody's variable compensation depends on, attached to a problem nobody's function owns.
What the answer usually looks like
The point of the two numbers is that a firm computes them for itself. What follows is the pattern I would expect to find, offered as a hypothesis to be tested rather than a conclusion to be adopted.
The processes with the highest cost to run are usually the high-volume transactional ones, Order to Cash and Procure to Pay. Their tail risk is genuinely low, because a bad call gets unwound that afternoon. Their accreting friction is not low, and is routinely mistaken for low, because each individual failure is trivial and the aggregate is how a firm teaches its customers that dealing with it is exhausting.
The processes with the highest tail risk are usually the small ones, Record to Report and Hire to Retire, which run on a handful of people and look like rounding errors on any cost-led roadmap. Small teams. Large tails.
Trouble to Resolve is usually high on both, and it carries the most accreting friction of any process in the firm, because it is where customers arrive already unhappy.
What would disconfirm this: a firm that computes both numbers honestly, including accreting friction, and finds its cost ranking and its value-at-risk ranking broadly agree. That firm should fund on cost and ignore the rest of this paper.
What to do with it
Fund where value at risk is highest, not where cost is highest, and expect the business case to be argued in two different rooms.
Tail risk is an audit and risk committee argument. What is being bought is the absence of an event, the payback cannot be modelled, and the transformation budget will always prefer a process with a payback number. Taking it through the transformation budget is how these programmes get refused.
Accreting friction is a CFO argument about the durability of the revenue line, and it is the stronger case commercially because it is not a hypothetical. The erosion is already happening. The question is only whether anyone has measured it.
Presenting the second argument as the first is the most common way a correct recommendation loses.
How much autonomy, and when
Which process gets money first and which gets to run without a human are different questions. Answering them the same way is how good programmes fail.
Flagging is the safest starting point, provided false-positive load is controlled. Letting AI tell a person a case looks unusual costs little when it is wrong, and it produces the record the system learns from. For most firms it will also be the first time the exception path has been measured at all, which makes it the cheapest route to the second number. Flag too liberally and the tier stops reading the flags, which is worse than not flagging.
Genuine novelty is never a candidate for autonomy, whatever the process. A case with no precedent is where a model is least reliable and most confident, which is the worst pairing available. Autonomy is earned only over the recurring-but-unspecified category, one process at a time, on four tests: how reversible the mistake is, how quickly it becomes visible, how far it spreads before anyone notices, and whether it carries regulatory consequence. A reconciliation entry caught at close scores well on the first two. A misjudged service-level dispute reaching a customer who was already angry scores badly on all four. Run the tests per process rather than assuming a universal order, and note that an autonomous accounting error clearing the close undetected fails them as badly as anything in the contact centre.
A bad forecast leaves a firm with staff it did not need. A bad autonomous decision creates the exact problem it was bought to prevent, in front of the worst possible audience.
One caution carries over from the individual layer, and the three categories sharpen it. Automating the routine erodes the operator's grip on the exception, because the routine was where the grip was formed. Take the recurring-but-unspecified cases into software and the escalation tier is left with genuine novelty only: the hardest diet available, on the thinnest pattern base, held by people who no longer see the forty-times-a-year cases that taught them the shape of the process. That tier gets more important as it gets less practised. The mitigation is deliberate exposure, not hope.
What to watch
| Signal | What it tells you | When to act |
|---|---|---|
| Share of cases still needing a human, by process | Whether recovery is compounding or has stalled | Two months flat |
| Cost of a case the system resolved against one it escalated | Whether the new path is genuinely cheaper than the old | The gap closes |
| Discount rate on renewals, split by exception-path exposure | The only direct read on accreting friction before the stall | The gap between the two groups widens |
| How often an autonomous decision gets reversed | Whether this process has earned more autonomy, or less | Above the threshold set for that process |
| Near-misses in the small, high-tail processes | The only early warning before the tail event itself | Any near-miss touching audit, law or regulation |
Review quarterly. Recompute value at risk annually, or after any near-miss.
Where the advantage will be
The temptation is to point AI at the largest cost lines, because that is where the money visibly is and where the payback arrives inside a fiscal year. Those programmes will work. They will also deliver the smallest strategic difference, because they are competing to automate the case every competitor has already automated twice.
The exception path is where the firm is least designed, least measured, least defended and most exposed. It is also now far cheaper to address than the operating model assumes.
AI will not create the biggest advantage where firms already spend the most. It will create it where firms have spent a century suppressing the exception rather than designing for it.
Appendix: the seven processes
What the exception looks like in each, and what AI does about it.
| Process | What the exception looks like | What AI does about it |
|---|---|---|
| Order to Cash | An invoice that will not clear: disputed terms, credit hold, tax edge case | Resolves it without escalation, and learns the pattern |
| Trouble to Resolve | A fault with no single root cause, or a disputed service level | Handles the ticket end to end; also forecasts the volume spike |
| Procure to Pay | A failed three-way match, or spend with no purchase order behind it | Clears the match, routes the approval, flags the off-contract pattern |
| Record to Report | A reconciliation break or a manual journal nobody can explain | Catches it before close rather than after publication |
| Hire to Retire | A leaver dispute, a comp exception, a cross-border employment question | Handles routine cases; stops short of anything legally contested |
| Forecast to Plan | An assumption two business units cannot reconcile | Reads more signals sooner, and surfaces the conflict earlier |
| Concept to Market | A decision taken on a picture of the market that was already stale | Market sensing, competitor monitoring, requirements synthesis, pricing |
On Concept to Market specifically: the intuitive reading is that it resists AI, because every product is bespoke and bespoke work has no pattern to learn. The creative act is bespoke. The apparatus around it is not. Market sensing, competitor monitoring, requirements synthesis, pricing analysis and portfolio optimisation are patterned, data-rich, and already among the fastest-moving areas of enterprise AI. What differs is the shape of the exception: it produces no error message and no queue, and surfaces two years later as a product nobody wanted. The exceptions that generate no exception report are the expensive ones.
Basis
Three related essays at jcorp.uk carry mechanisms this paper depends on: The internet in the rain (2023) on service levels as negotiated settlement; Customer emotional debt (2023) on debt that accumulates behind a dam and comes due as lost pricing power; The clerk we fired (2026) on displaced cognition at the individual layer.
This paper contains no proprietary data and asserts no measured positions for the seven processes. The pattern described under "What the answer usually looks like" is a hypothesis derived from the structural argument, and it is offered to be tested against a firm's own two numbers rather than adopted.
- Central claim
- AI makes judgment cheap for the first time, so firms should fund AI by value at risk, the exception path's tail risk and accreting friction, not by cost to run on the standard case that two prior waves of automation already made cheap.
- Upstream variable
- Value at risk, the variable one level upstream of cost to run
- Concepts
- The variable one level upstream
- Related
- The clerk we firedThe internet in the rainCustomer emotional debt: the strategic risk you're not forecastingDelivering value creation