Almost every process automation pitch starts from the same question: can AI do this? It is the wrong question. The useful one is narrower. Which part of this process can AI do repeatedly, measurably and reversibly, at a risk-adjusted cost below the human alternative, and how will we know when a case falls outside that part?
The evidence for AI in operations is genuinely strong, and considerably more specific than the marketing suggests.
The numbers that hold up
Controlled studies, as opposed to vendor case studies, converge on a consistent picture for bounded, repetitive work.
Operational deployments report larger numbers again: around 70% faster invoice processing at Careem, 20 to 50% lower forecasting error, over 30% fewer defects on suitable manufacturing lines. Those are real outcomes, but each describes a redesigned system, with clean data, integration, business rules, monitoring and human exception handling. A model dropped into an existing workflow produces none of that.
The finding that should change how you scope
In an experiment with 758 BCG consultants, AI users completed tasks roughly 25% faster with quality rated about 40% higher, on tasks inside the model's capability frontier. On a problem just outside it, AI-assisted consultants were 19 percentage points less likely to reach the correct answer.
An AI system can be excellent on yesterday's benchmark and confidently wrong on the unusual case that actually matters.
This is the central scoping problem. Average accuracy tells you nothing about behaviour at the edge, and the edge is usually where the expensive cases live. It is why we treat "the model performs well" and "this process can be automated" as two entirely separate claims.
Three modes, not two
Most automation conversations collapse into automate or don't. Three modes are more useful, because "don't automate" rarely means "don't use AI". It means the system should not hold unilateral authority over the outcome.
Automate
Errors are cheap or reversible, outputs are objectively testable, inputs are stable, and exceptions can be detected automatically.
- Invoice extraction and matching
- Ticket classification and routing
- Demand forecasting
- Visual defect inspection
- Internal knowledge retrieval with citations
Augment
AI accelerates analysis or drafting, but a knowledgeable human still owns the decision and can be held to it.
- Contract review and clause extraction
- Recruiting screening, never the decision
- Lead prioritisation
- Analysis, research and first drafts
- Clinical and legal interpretation
Keep human-led
The decision affects someone's rights, health, livelihood, safety or access to an essential service. Or the situation is novel rather than repetitive.
- Hiring, firing, promotion, pay
- Credit denial and eligibility
- Diagnosis and treatment
- Safety-critical shutdown
- High-value or unusual payment release
The shape of a good candidate
Strong candidates nearly always follow one pattern: receive a standardised input, classify or extract or predict, check confidence against business rules, take a reversible action, escalate the exceptions, record everything. The further a workflow drifts from that shape, toward ambiguity, conflicting goals, negotiation, empathy or irreversible consequences, the stronger the case for augmentation instead.
Before committing, we score a process on ten factors: ROI after review costs, frequency, volume, complexity, input variability, regulatory exposure, how much human judgement is central, explainability, data availability, and error cost. A high total never overrides a veto: irreversible safety or fundamental-rights risk stops the conversation regardless of the score.
Realistic ranges by process
| Process | Mode | Realistic gain | What actually bites |
|---|---|---|---|
| Customer service | Automate tier 0 and 1 | 15 to 35% agent productivity | Hallucinated policy, escalation failure |
| Accounts payable | Automate with controls | 40 to 70% processing time | Fraud, duplicate payments, approval authority |
| Recruiting | Augment only | 20 to 40% on admin, 0% on selection | Discrimination, no validated predictive lift |
| Legal review | Augment | 50 to 80% on structured extraction | Hallucinated authority, privilege |
| Software and analysis | Augment | 20 to 40%, up to 55% bounded | Fabricated evidence, wrong outside frontier |
| Forecasting | Automate | 20 to 50% lower forecast error | Regime changes, bad master data |
What the ROI numbers hide
Deloitte surveyed 1,854 executives and found satisfactory ROI on an AI use case typically took two to four years, with only 6% seeing payback inside a year. McKinsey, looking at operational leaders, found six to twelve months. Both are true. Contained operational automation on mature data pays back fast; enterprise-wide AI transformation does not.
The term most business cases omit entirely is expected error loss. A process that saves a million in labour but carries a 0.1% chance of a twenty-million regulatory or safety event does not have a million-dollar business case. Average accuracy does not settle it either, because subgroup errors, false negatives and behaviour under distribution shift usually matter more than the headline score.
- The denominator problem: automating the easiest 60% of tickets may remove far less than 60% of the labour, because the remaining cases consume more minutes each
- The attribution problem: most published wins bundle AI with CRM migration, data cleanup, retraining and process redesign
- The silent labour problem: review work that moves into a dashboard is still work, and it belongs in the denominator
"Human in the loop" is not a safety mechanism
The most common control we see is also the weakest as usually implemented. In a pathology experiment, AI raised average performance but still produced a 7% automation-bias rate: cases where a correct human judgement was reversed after bad AI advice. A separate randomised study of 2,784 participants found that making review procedures more burdensome increased acceptance of incorrect suggestions.
An Approve button is not oversight
Real review means the human sees the underlying evidence rather than just the recommendation, has time and authority to disagree, sometimes records their own assessment before the model's is revealed, and is periodically tested on whether they actually catch model errors. Anything less produces the appearance of accountability while preserving the bias.
Where we start
The most rational portfolio starts with boring processes. High-volume extraction, classification, forecasting, comparison, inspection, drafting and reconciliation produce better risk-adjusted economics than autonomous agents making consequential decisions. They are also how you build the monitoring and exception-handling muscle you will need later anyway.
Then set the error threshold from business harm backwards, not from whichever benchmark a vendor supplies. A 95%-accurate system with instant detection and free reversal is a far safer automation than a 99.9%-accurate one whose rare mistakes are irreversible.
That is the work our forward deployed engineers do before writing any code: map the process, find the subset that survives this test, and say plainly when the honest answer is conventional automation, a process fix, or leaving it alone.