Which business processes are worth automating with AI, and which aren't

Fadel Dia-Eddine· Co-Founder & Product Lead6 min read

Almost every process automation pitch starts from the same question: can AI do this? It is the wrong question. The useful one is narrower. Which part of this process can AI do repeatedly, measurably and reversibly, at a risk-adjusted cost below the human alternative, and how will we know when a case falls outside that part?

The evidence for AI in operations is genuinely strong, and considerably more specific than the marketing suggests.

The numbers that hold up

Controlled studies, as opposed to vendor case studies, converge on a consistent picture for bounded, repetitive work.

14%
Customer support productivity
NBER, generative AI assistant
35%
Gain for less-experienced agents
Same study: AI transfers tacit knowledge
55%
Faster on a bounded coding task
GitHub randomised trial
40%
Faster on professional writing
Noy & Zhang

Operational deployments report larger numbers again: around 70% faster invoice processing at Careem, 20 to 50% lower forecasting error, over 30% fewer defects on suitable manufacturing lines. Those are real outcomes, but each describes a redesigned system, with clean data, integration, business rules, monitoring and human exception handling. A model dropped into an existing workflow produces none of that.

The finding that should change how you scope

In an experiment with 758 BCG consultants, AI users completed tasks roughly 25% faster with quality rated about 40% higher, on tasks inside the model's capability frontier. On a problem just outside it, AI-assisted consultants were 19 percentage points less likely to reach the correct answer.

An AI system can be excellent on yesterday's benchmark and confidently wrong on the unusual case that actually matters.

This is the central scoping problem. Average accuracy tells you nothing about behaviour at the edge, and the edge is usually where the expensive cases live. It is why we treat "the model performs well" and "this process can be automated" as two entirely separate claims.

Three modes, not two

Most automation conversations collapse into automate or don't. Three modes are more useful, because "don't automate" rarely means "don't use AI". It means the system should not hold unilateral authority over the outcome.

Automate

Errors are cheap or reversible, outputs are objectively testable, inputs are stable, and exceptions can be detected automatically.

  • Invoice extraction and matching
  • Ticket classification and routing
  • Demand forecasting
  • Visual defect inspection
  • Internal knowledge retrieval with citations

Augment

AI accelerates analysis or drafting, but a knowledgeable human still owns the decision and can be held to it.

  • Contract review and clause extraction
  • Recruiting screening, never the decision
  • Lead prioritisation
  • Analysis, research and first drafts
  • Clinical and legal interpretation

Keep human-led

The decision affects someone's rights, health, livelihood, safety or access to an essential service. Or the situation is novel rather than repetitive.

  • Hiring, firing, promotion, pay
  • Credit denial and eligibility
  • Diagnosis and treatment
  • Safety-critical shutdown
  • High-value or unusual payment release

The shape of a good candidate

Strong candidates nearly always follow one pattern: receive a standardised input, classify or extract or predict, check confidence against business rules, take a reversible action, escalate the exceptions, record everything. The further a workflow drifts from that shape, toward ambiguity, conflicting goals, negotiation, empathy or irreversible consequences, the stronger the case for augmentation instead.

Before committing, we score a process on ten factors: ROI after review costs, frequency, volume, complexity, input variability, regulatory exposure, how much human judgement is central, explainability, data availability, and error cost. A high total never overrides a veto: irreversible safety or fundamental-rights risk stops the conversation regardless of the score.

Realistic ranges by process

ProcessModeRealistic gainWhat actually bites
Customer serviceAutomate tier 0 and 115 to 35% agent productivityHallucinated policy, escalation failure
Accounts payableAutomate with controls40 to 70% processing timeFraud, duplicate payments, approval authority
RecruitingAugment only20 to 40% on admin, 0% on selectionDiscrimination, no validated predictive lift
Legal reviewAugment50 to 80% on structured extractionHallucinated authority, privilege
Software and analysisAugment20 to 40%, up to 55% boundedFabricated evidence, wrong outside frontier
ForecastingAutomate20 to 50% lower forecast errorRegime changes, bad master data
Planning ranges, not guarantees. Published studies vary widely in task definition and implementation maturity.

What the ROI numbers hide

Deloitte surveyed 1,854 executives and found satisfactory ROI on an AI use case typically took two to four years, with only 6% seeing payback inside a year. McKinsey, looking at operational leaders, found six to twelve months. Both are true. Contained operational automation on mature data pays back fast; enterprise-wide AI transformation does not.

The term most business cases omit entirely is expected error loss. A process that saves a million in labour but carries a 0.1% chance of a twenty-million regulatory or safety event does not have a million-dollar business case. Average accuracy does not settle it either, because subgroup errors, false negatives and behaviour under distribution shift usually matter more than the headline score.

  • The denominator problem: automating the easiest 60% of tickets may remove far less than 60% of the labour, because the remaining cases consume more minutes each
  • The attribution problem: most published wins bundle AI with CRM migration, data cleanup, retraining and process redesign
  • The silent labour problem: review work that moves into a dashboard is still work, and it belongs in the denominator

"Human in the loop" is not a safety mechanism

The most common control we see is also the weakest as usually implemented. In a pathology experiment, AI raised average performance but still produced a 7% automation-bias rate: cases where a correct human judgement was reversed after bad AI advice. A separate randomised study of 2,784 participants found that making review procedures more burdensome increased acceptance of incorrect suggestions.

An Approve button is not oversight

Real review means the human sees the underlying evidence rather than just the recommendation, has time and authority to disagree, sometimes records their own assessment before the model's is revealed, and is periodically tested on whether they actually catch model errors. Anything less produces the appearance of accountability while preserving the bias.

Where we start

The most rational portfolio starts with boring processes. High-volume extraction, classification, forecasting, comparison, inspection, drafting and reconciliation produce better risk-adjusted economics than autonomous agents making consequential decisions. They are also how you build the monitoring and exception-handling muscle you will need later anyway.

Then set the error threshold from business harm backwards, not from whichever benchmark a vendor supplies. A 95%-accurate system with instant detection and free reversal is a far safer automation than a 99.9%-accurate one whose rare mistakes are irreversible.

That is the work our forward deployed engineers do before writing any code: map the process, find the subset that survives this test, and say plainly when the honest answer is conventional automation, a process fix, or leaving it alone.

AIProcess automation

Questions, answered

Score it on frequency, volume, how bounded the task is, input variability, regulatory exposure, how central human judgement is, explainability, data availability, and error cost. Then check whether the benefit survives review, monitoring and governance costs. If an error would affect someone's rights, health, livelihood or access to an essential service, treat it as augmentation regardless of the score.

For a contained operational workflow on mature data (invoice processing, ticket classification, forecasting) six to twelve months is realistic. For anything requiring new data platforms, operating-model change or organisational redesign, Deloitte's survey of 1,854 executives found two to four years is more typical, with only 6% seeing payback inside a year.

Not on its own. Automation bias means reviewers accept incorrect AI recommendations at measurable rates. One pathology experiment found 7% of correct human judgements were reversed after bad AI advice. Effective oversight requires the reviewer to see the underlying evidence, have time and authority to disagree, and be tested periodically on whether they actually detect model errors.

Decisions affecting someone's livelihood, health, safety, liberty or access to essential services: hiring and firing, medical diagnosis and treatment, legal advice and filings, credit denial, disciplinary investigations, safety-critical shutdown, and release of high-value or unusual payments. AI can analyse and recommend in all of these. An accountable human should decide.

Working on something similar?

Alpine Edge builds and runs this kind of system for clients across Europe and MENA. Tell us what you are trying to solve and we will tell you how we would approach it.

Talk to an engineer

Read next

All articles