The market spent this year selling AI by the job title. Two years of real deployments are now in, and they divide along a line the category did not draw: whether the work reconciles against a system of record, and whether escalation was designed in.
The market spent this year selling artificial intelligence by the job title. An AI sales representative. An AI recruiter. An AI accountant. An AI support agent. Enough of those deployments have now run long enough to be judged, and they divide cleanly into the ones that held and the ones that were quietly undone. The line between them is not model quality, and it is not how difficult the job looked.
It is worth being precise about what happened, because the failures were not quiet and the successes were not the ones the category was marketed on. A company that replaced 700 support agents is hiring them back. A company that replaced its bookkeepers never looked up. Both used competent models. The difference lies in the work itself, and in what the business built around it.
Independent research puts the AI workforce market at roughly $37 billion in 2025, growing toward a projected $563 billion by 2035. Customer support is the largest single function at about a quarter of the market. Capital agrees: global venture funding reached $510 billion in the first half of 2026, more than all of 2025 combined, with AI absorbing more than seventy percent of second-quarter investment. The top twenty-five agent companies have raised over $25 billion between them.
The commercial model moved with it. Vendors stopped selling seats and started selling finished work. Sierra, valued at $15.8 billion, charges per resolved conversation and nothing at all for the conversations it escalates to a person. That is not a pricing experiment. It is the category conceding that if software is doing the job, the customer should pay for the job.
None of this is in dispute. What is in dispute is whether the thing being bought survives contact with the business that bought it.
| Company | What was replaced | Where it stands |
|---|---|---|
| Pilot | Bookkeepers, replaced by a proprietary AI accountant. | Sustained |
| IBM | Back-office roles across HR, finance and administration. | Sustained, with further reductions planned |
| Salesforce | Roughly 4,000 support roles, using its own agent platform. | Sustained |
| Wendy's | Drive-through order taking, reported at 86% accuracy. | Scaling past 500 locations |
| Klarna | The work of around 700 customer service agents. | Reversed; rehiring human agents |
| Artisan | Sales development, handled by an outbound AI agent. | Reverted to a hybrid model |
| McDonald's | Drive-through order taking. | Discontinued |
| Duolingo | Translation and content writing contractors. | Quality complaints; walked back |
The Klarna case is the most instructive because the company was unusually candid about it. Its assistant handled two-thirds of incoming queries. On volume, it worked. On quality, it did not: complaints accumulated, and they concentrated in refunds and billing, which is precisely the territory where a confident wrong answer costs more than no answer at all. The chief executive's summary of what went wrong was one sentence long.
We focused too much on cost.
Sebastian Siemiatkowski, chief executive, Klarna
The individual reversals are not outliers. A 2026 study drawing on 52 executive interviews, surveys of 153 leaders and an analysis of 300 public deployments found that 95% of enterprise pilots delivered no measurable impact on profit and loss. The attrition happens earlier than most buyers expect. Sixty percent of firms evaluated an enterprise-grade system. Twenty percent got as far as a pilot. Five percent went live.
Figure 1: how far enterprise AI initiatives actually get
The researchers were direct about the cause: the failures were organizational rather than technical. The deployments that survived shared three properties. They were integrated into the core workflow rather than bolted alongside it, they retained memory across interactions, and they adapted to feedback. Nothing on that list is a model capability.
Plot the real deployments against two questions and the pattern resolves. First: is there a system of record the output can be reconciled against before it becomes a business fact? Second: was a path to a human designed in from the start, with the cost of that oversight in the budget?
Figure 2: real deployments plotted against ground truth and designed escalation
Bookkeeping is not a simple job, but a posted entry either reconciles or it does not. A drive-through order appears on a screen the customer confirms before anything is cooked. A back-office record settles against a system that already holds the answer. In each case the business can tell whether the work was done correctly at the moment it is done, without waiting for a customer to complain.
A refund decision has no such check. Neither does a cold outreach email or a translated paragraph. The output is a judgment, it reaches a person immediately, and the only feedback signal is dissatisfaction arriving weeks later in aggregate. That is the quadrant every reversal came from.
An independent teardown of six commercial AI employee platforms reached a finding that should reshape how these systems are costed. Human oversight consumes between 30% and 50% of total AI investment in practice. The savings are real, but they are roughly half the number on the slide, and the missing half does not appear as a line item until the deployment is already running.
The same review identified why the burden stays so high, and the diagnosis is architectural rather than commercial.
The agents work in isolation.
Each one holds its own context. Nothing is shared between them, so the coordination a human team performs without thinking has to be performed by a person instead.
The data integration is shallow.
Most platforms generate text convincingly and reach into systems of record barely at all, which is exactly the capability that would have supplied the ground truth in Figure 2.
Oversight is unstructured.
Where there is no defined approval boundary, review becomes a person reading output and forming an opinion. That work is unbounded, unpriced and impossible to reduce.
These three failures are the same failure. An AI employee sold as a role, disconnected from the company's systems and governed by nothing in particular, cannot reconcile its own work. So a person reconciles it, indefinitely, and the economics that justified the purchase never arrive.
Each one is a product decision rather than a law of nature, and AIOS made the opposite decision in all three cases. Against isolation, one cognitive core that carries knowledge, memory and policy across every task rather than a fleet of agents that each hold their own. Against shallow integration, a shared connector fabric so execution lands in the systems of record that already hold the answer. Against unstructured oversight, an approval boundary declared in the workflow, with a defined owner and a known cost. Those are the two axes of Figure 2, supplied by the platform instead of hoped for in each deployment.
A job title is a bundle of tasks a company assembled around one person's working hours. It is a convenient hiring abstraction and a poor automation boundary, because the tasks inside it have wildly different verifiability and wildly different consequences for being wrong. Buy the bundle and you have bought the reconcilable work and the unreconcilable judgment as one indivisible unit, priced as though both were safe.
AIOS is built on the opposite unit. The thing that gets automated is a task with an agreed definition of done, executed by one cognitive core against the systems of record through a shared connector fabric, with the approval boundary declared in the workflow rather than improvised by whoever is watching. Ground truth is not a property the task happens to have. It is something the platform supplies, because the execution runs through the systems that hold the answer.
Escalation is handled the same way. A human decision point is a defined step with a cost and an owner, not an admission that the automation fell short. That is what converts the 30% to 50% oversight tax from an unbounded drain into a priced, bounded part of the work.
It is also what makes outcome-based pricing honest rather than aspirational. A platform can only charge for finished work if it can prove the work finished, and it can only prove that if execution touched a system of record and left evidence behind. We wrote about that mechanism separately in Pricing what you can prove.
The 95% figure measures one specific outcome: work that ran, produced something, and never became part of how the business operates. Three conditions have to hold for a project to end that way. An AIOS engagement closes all three before anyone builds anything, which is a different kind of claim than a good track record.
Nobody agreed in advance what would count.
A pilot with no baseline is graded afterwards against a standard nobody wrote down, which is how a project produces no measurable impact on profit and loss without anyone being able to say when it went wrong. We baseline against what the process costs and produces today, and both sides sign that before work starts. The measure exists before the work does, so success cannot be redefined at the end and failure cannot hide.
The work had no path into production.
Most pilots are built beside the business and have to be re-implemented to enter it, and that second build is where they die. In AIOS there is no second build. The workflow proved in the pilot is the workflow that runs in production: same definition, same connectors, same policy, same approvals, same execution ledger. Going live is a promotion, not a rewrite.
Someone could get paid anyway.
This is the condition the industry avoids discussing. Under per-seat licensing a vendor is paid whether or not the work ever happens, so shelfware is commercially survivable. Under outcome pricing it is not. Fees attach to completed tasks, failed tasks are free, and a workflow that never reaches production completes nothing and bills nothing. There is no revenue model here for a project that stalls.
Read those together and the reason work reaches production stops being a matter of confidence. It is closed by the engagement, by the architecture and by the commercial model respectively. A stranded pilot is not a risk we manage. It is an outcome the model does not pay for.
The lesson of the reversals is not that AI cannot do enterprise work. Klarna's assistant handled two-thirds of its queries, and it was pointed at the third that needed judgment anyway. The lesson is that the job title was the wrong thing to buy. Automate the task you can prove, escalate the judgment you cannot, and put the boundary between them in the system rather than in someone's discretion.
Sources
We embed until it works, then you pay for what worked. Bring the process you would most like to stop staffing.