AI pilots are no longer hard to impress people with. A model can read a document, classify a request, compare supplier quotations, draft a reply, or flag a pattern that a busy team would have found later. In a controlled demonstration, that is often enough. In a live operation, it rarely is.
Production work does not arrive as a clean sample. An invoice is missing a purchase-order number. Two systems disagree about a supplier record. A field report leaves out the one figure a supervisor needs. An approval sits with someone who is traveling. A policy exception is handled, today, by a phone call that never enters the system. The model may still be capable. The operation around it is not ready to use that capability without creating new work, new risk, or both.
That is the failure most businesses miss. They ask whether AI can perform the task. The more useful question is how the operation should work, and which stations inside that operation AI should be allowed to run.
The readiness gap is already visible in enterprise research. In a 2026 Deloitte survey of 501 U.S. leaders involved in agentic AI, only 5 percent said their business processes were highly prepared for AI agents. At the same time, 74 percent expected nearly half of their processes to be redesigned or rebuilt around agents within four years. Gartner has separately predicted that more than 40 percent of agentic AI projects will be canceled by the end of 2027, citing escalating cost, unclear business value, and inadequate risk controls. The pattern in both findings is the same. Ambition is ahead of the workflow.
A pilot proves a capability. Production requires a system.
A pilot is usually built to answer a narrow question. Can the model extract the right fields from these invoices? Can it classify these support requests? Can it summarize these contracts? Those are fair tests, and a good result is worth having. They do not test the operation.
Once the same capability is asked to run on live work, a different set of questions appears. Where does the input come from, and who is allowed to submit it? Which system holds the authoritative record when two sources conflict? What may the system do on its own, and what must wait for a person? What happens when required information is missing, when confidence is low, or when the reviewer does not respond? What is written down so that someone can reconstruct the decision six months later?
If those questions are unanswered, the pilot has not failed because the model was weak. It has failed because it was inserted into a process that still depends on informal knowledge. Different teams follow different rules. Exceptions move through chat. The approval path lives in someone's head. AI added to that kind of process does not stabilize it. It makes the instability faster and harder to see.
Automate the station, not the job title
A practical way through this is to stop designing around roles. "An AI accountant" or "an AI procurement officer" is not a system. It is a wish. A finance officer's day is a sequence of different stations: reading invoices, matching them to purchase records, resolving mismatches, preparing entries, and making judgments about what is material. Those stations do not have the same risk, the same rules, or the same need for a person.
The useful design is smaller. Take accounts payable. The weak starting point is a general assistant for the finance team. A workable starting point is a defined station:
An incoming supplier invoice is received through a controlled channel. The system reads the document, extracts the fields the process actually uses, and checks them against the purchase record and goods receipt. Matches within defined tolerances are prepared for posting. Mismatches, missing references, duplicate invoice numbers, and amounts above a threshold are assembled into an exception pack for a named finance officer. The officer reviews the exception, not the entire pile. The decision, the source documents, and the outcome are written back to the finance record.
That station does real work. It also has a boundary. It is not being asked to own accounting judgment, revenue recognition, or the relationship with the supplier. The same cut can be made in procurement, case handling, field reporting, and compliance review. AI takes the repeatable interpretive station. A person keeps the station where authority, exception, or consequence sits.
What belongs in the workflow, not in the prompt
A reliable station has five parts, and only one of them is the model.
Start with observation. Map what actually enters the process: an email, a portal form, a PDF, a system event, a field submission. Note what a person does with it today, including the checks they make from memory. Do not begin with model features. Begin with the work, including the work that is currently invisible because it happens outside the system.
Then define interpretation. This is where current models earn their place: reading unstructured documents, classifying free text, comparing a submission with a record, extracting the fields a downstream system needs, and summarizing a case for a reviewer. Be specific about the sources. If the authoritative supplier status lives in one system and the quotation lives in another, the station needs both, or it will reason fluently over an incomplete picture.
Next, define the action. "Help with procurement" cannot be tested, permissioned, or audited. "After the requisition is complete, retrieve approved-supplier status and the last three purchases, then prepare a comparison for the buyer" can. Typical actions at this stage are narrow: create a draft record, update a status, generate a document, notify the next owner, or place a standard case on the normal path. The action should be something the organization already knows how to reverse or review.
Escalation is part of the design, not a fallback added after go-live. The station should stop and hand work to a named role when required information is missing, two records conflict, the value crosses a threshold, confidence is below the agreed level, a policy exception appears, or the action would have material financial, customer, or compliance impact. Also define the timeout. If the reviewer does not respond, the work should not vanish into a queue nobody owns.
Finally, record the run. What came in. What the system extracted or decided. Which rule or threshold fired. Who approved the exception. What the outcome was. This is what makes a later dispute, audit, or improvement cycle possible. It is also what separates a workflow from a chat session that nobody can reconstruct.
Current agent platforms are starting to reflect the same shape, including flows that pause for a person before continuing. The tooling is not the point. A human checkpoint that does not say who reviews, what they are looking for, and what authority they have is only a pause.

One workflow, before and after
Consider a procurement team that wants AI because quotation comparison is slow. In the current process, a department emails a requisition. Someone checks, informally, whether the specification and budget code are present. Procurement searches old files for supplier history, chases quotations, builds a comparison in a spreadsheet, sends it to a manager, waits, then raises a purchase order and tells the supplier. A pilot that compares three clean quotations can look successful and still leave almost all of that work untouched.
A production design starts further back. The requisition enters through one channel, with the fields the process requires. The system checks completeness and returns missing items to the requester instead of letting a half-specified request occupy a buyer. If the request is complete, it pulls approved-supplier status and recent purchase history, organizes the quotations it is allowed to use, and prepares the comparison. Standard purchases inside defined limits follow the company's rule. High value, a new supplier, conflicting records, or a policy exception go to the buyer or the approving manager, with the comparison and the reason for escalation attached. After approval, the purchase order is prepared, the record is updated, the supplier is notified, and the decision history stays with the transaction.
The comparison model is still in the workflow. It is no longer the workflow. Most of the operational gain comes from closing the gaps around it: one intake, a completeness check, a known source of supplier truth, a rule for the normal path, and a person on the exception path.
The same cut applies elsewhere, with different stations and different boundaries. In customer operations, AI can classify a request, retrieve account context, resolve a standard low-risk case, and route the rest; a person keeps complaints, credits, and anything that changes the customer relationship. In field operations, AI can check submitted forms, list missing items, and prepare the supervisor's follow-up list; the supervisor still interprets unusual site conditions. In compliance review, AI can test completeness, extract required fields, and brief the reviewer; the authority to accept or reject stays with the reviewer. In management reporting, AI can assemble approved figures and draft the movement commentary; management still decides what the movement means.
Why the demo does not survive contact with the operation
Five problems show up repeatedly once a pilot is asked to run on live work.
The surrounding process is unclear. The AI task was specified. The operation was not. Teams disagree about what "complete" means, approvals depend on who is available, and exceptions are resolved in side channels. Automating a station inside an unstable process encodes the instability.
The data the station needs is fragmented. Pilots are fed selected examples. Production depends on email, shared drives, spreadsheets, a legacy system, and a field someone maintains locally. A model can reason over whatever it is given. It cannot invent the authoritative record it was never connected to. If two sources disagree, the workflow needs a rule for which one wins, or a path that stops and asks.
The normal case works and the exception does not. Demonstrations follow the expected path. Live work includes duplicate invoices, incomplete supplier records, unusual customer requests, and policies that changed last month. A production station needs an explicit outcome for each class of break: return to the requester, escalate to a role, or refuse to act. "The model will figure it out" is not an exception path.
Human responsibility is named too vaguely. "Human in the loop" does not say which human, at which point, reviewing what, with what authority, or what happens if they do not respond. Deloitte's 2026 work on agent orchestration describes the practical range as humans in the loop, on the loop, or outside the loop, chosen according to task complexity and consequence. That choice has to be made per station. A low-risk status update and a payment release are not the same decision.
Success is measured on the model instead of the operation. Extraction accuracy and answer quality matter, and they should be tracked. They do not tell you whether the backlog fell, whether requests stopped getting lost, whether reviewers spent less time on clean cases, whether exception cycle time improved, or whether cost per transaction moved. A pilot can score well on sample documents and still leave the operation unchanged.
Boundaries are what make a stronger station possible
Controls are sometimes treated as a way of making AI smaller. In an operation, they are what make a larger role defensible. If the station has a defined input, a permitted action, a threshold for stopping, and a person accountable for the exception, the organization can let it handle more of the normal path without pretending it has judgment it does not have.
The right level of autonomy is a property of the task, not a philosophy. Preparing a standard acknowledgment, or matching an invoice inside tolerance, can run with little intervention once the rule is stable. Releasing a payment, accepting a compliance exception, or answering a complaint that affects the relationship should stop for a person even if the draft is good. Both can sit in the same operation. What does not work is one autonomy setting applied to every station because the project was framed as "deploying an agent."
Start with one loop you can measure
An organization does not need to redesign a department before it can use AI. It needs one loop where people repeatedly read something, classify or compare it, move it between systems, and prepare a standard output. The test is plain. Can the interpretation be checked against a known set of cases? Can the next action be stated as a rule? Can the exceptions be recognized before they reach a customer or a ledger? Is there a named owner for the queue? Can the team see, after a month, whether cycle time, rework, or backlog actually moved?
If those answers are yes, run that loop until the exception list is boring. Then extend it. Expanding from a station that already records its failures is much safer than launching a second pilot because the first demonstration looked impressive.
AI does not need to replace a function to change how an operation runs. It needs a station with an input, a permitted action, a reason to stop, and a record of what it did. The organizations that get past the pilot are usually the ones that designed that station before they asked the model to perform it.
Sources
Deloitte, "AI Agents are Only the Beginning," press release, 12 August 2026. Survey of 501 U.S. senior manager to C-suite respondents involved in agentic AI. Only 5 percent said business processes were highly prepared for AI agents; 74 percent expected nearly half of processes to be redesigned or rebuilt around AI agents within four years. https://www.deloitte.com/us/en/about/press-room/deloitte-survey-examines-ai-readiness-agentic-ai-success.html
Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," 25 June 2025. Cancellation drivers cited: escalating costs, unclear business value, inadequate risk controls. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
Deloitte, "Unlocking exponential value with AI agent orchestration," TMT Predictions 2026. Autonomy described as a spectrum: humans in the loop, on the loop, or out of the loop, depending on task complexity, domain, workflow design, and outcome criticality. https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/ai-agent-orchestration.html