AI Agents can finally do the work. But at enterprises, it'll take a while.
Why the work that AI can finally do is the work nobody can explain
I was talking to a SVP at an enterprise recently, who manages a pretty big operations team. He generally understands what the team accomplishes and that it’s important, but is hazy on the details. He knows that the work is something ChatGPT could do. He would love to have an agent do the work - he’d be a hero with all the money he’d save (the company is in a very cost pressured industry). He should be able to make it happen - they all report to him.
It’s not going to happen. And the reason has nothing to do with the model.
In 2014, the economist David Autor made a point that felt obvious at the time: the work that’s hard to explain is the work that’s hard to automate. We can do all sorts of things we can’t put into words, and you can’t hand a machine a task you can’t describe. Robots install windshields on the assembly line, but aftermarket replacement is done by humans. The line between what could and couldn’t be automated was drawn by what a machine could understand.
That line just moved. LLMs are the first technology that can handle ambiguity, absorb context, and act in a messy situation. So is there any boundary left? If not, why does almost everyone still have a job?
The trouble is that our SVP doesn’t really understand how their teams do the actual work. Neither does the VP who reports to them. Nor the Director. The manager used to do the work, so has some idea, but hasn’t seen the systems in a while.
And that means that the team is not going anywhere.
The SVP will ask for the automation. And everyone will nod their heads and agree. The VP, and director also want the cost savings, the manager and the folks doing the work want to look like team players, and they will be excellent ones. The project will start, someone will ask for requirements, and the answers will be presented by the managers and the team. They will be complete, accurate, and useless. And they’ll be delayed. And the person with the answer at each stage will be on vacation a surprising amount of the time. And emails will land in spam a surprising amount of the time.
The project is doomed.
Nobody has to refuse the automations. We don’t need Luddites going out and breaking machines. It just doesn’t happen, because the people who have to be tasked with making it happen don’t want it to happen. No sabotage necessary.
Contact centers fell at the front door
Contact centers are one place where this pattern has broken and AI has made a meaningful difference. Since some of the tickets reaching contact centers can be classified and siphoned off to the desired result, the team doing the work just doesn’t see them. As long as the result is known, and the work is not complex, the AI can do it. It can handle a password reset, tell you about the order status, or about store hours.
But the 11 step process required to return an item that originally shipped from a drop shipper in Thailand via Vietnam? With steps 5c-5g undocumented? The AI isn’t doing that.
There were always two constraints
Until now, the constraint for automation was that the machine couldn’t do the work. It could do what it could do, and past that boundary it couldn’t. That boundary just became fuzzy. But in a lot of cases, it won’t matter, because there was always a second constraint that got conflated with the first one: nobody can tell the machine what the work to automate is. We’re about to experience a lot more of that.
Fine, watch the screen
Just have the AI watch what people are doing on their computers. Record the screen and the keystrokes. If you own the computers then you can do that, right?
I doubt it. First, try telling someone in America that their computer is going to be monitored so you can automate them. It won’t be pleasant. But also, I don’t think it’ll work.
Traces capture actions the user took. What they typed, what part of the screen they clicked on. But you don’t know the full extent of the judgment call they made. Or the things they glanced at and then ignored. Or the actions they didn’t take for a good reason. You’re not going to have the volume of data that it takes to infer that, so you’re stuck with partial information.
Which at least current models can’t make sense of anyway. The context of the job, the screen, and the action need to be provided to the LLM, and no one is going to be providing them.
Fine, watch the software
What if we combine the screen with what the software sees? The user is taking a bunch of actions, but at the same time they’re working in some SaaS software, which has its side of the story. Surely, combining the two gives a more complete picture.
I think the software has a lot of the missing context that the traces are missing, and the two combined would be pretty close to the whole picture. But consider the incentives of the software provider: their existence and lock in and margins are dependent on the person in the seat. They, too, don’t want to do anything to threaten that person’s job. The employee doing the 11 step process with the crazy workaround because the software was inefficient is now an unlikely ally of the very software company that made their life frustrating. The livelihoods of the two are suddenly intertwined.
There’s always a lineshaft
In 1990, an economic historian named Paul David tried to figure out why the advantages conferred by the computer weren’t showing up in the productivity statistics. He drew from an older analogy - the diffusion of the electric dynamo, which was available in 1880s. But in the productivity statistics, there was actually a productivity slowdown in Britain and the US from 1890 - 1913, the exact time period where the dynamo should have been improving productivity by replacing steam.
And as it turned out, the diffusion lag was very real. At the turn of the century, electric motors were under 5% of factory mechanical drive. Momentum only arrived in 1914 - 1917, when utility rates fell relative to prices and central stations overtook on-site generation. The delay itself was something else: plants were built for a world of steam, and even though no one was building any more plants for steam, it was unprofitable to replace the existing plants. So, the significantly superior technology of the dynamo would take decades more, and it was only in the 1920s capital boom that even half of factory drive got electrified. And the productivity statistics also followed.
The key part was the lineshaft. A single prime mover, a steam engine, made rotary power at one point in the plant. There wasn’t a way to move power except mechanically, so a long steel shaft ran the length of the whole ceiling, carrying that power down to every machine through pulleys and belts.
The shaft was what everything was built around, but that’s not how dynamos work. The first thing factories did was buy a dynamo and keep the shaft, using motors to turn sections of it. That’s group drive, and it bought them almost nothing. The gain came with unit drive: a motor on each machine, the shaft taken down, and the building laid out around the work instead of around the power. The productivity was never in the motor. It was in the layout the motor allowed.
Every company has one of these. The lineshaft is the process: the accumulated, undocumented way the work actually gets done, that nobody chose and nobody wrote down. The org chart and the software are the floor plan, arranged around it. And like the factories in 1900, you can buy the new thing, bolt it on, and change nothing, because the gain was never in the motor.
Yes it’s great, but see me in 20 years
The SVP has been very impressed with all the AI Agent demos. The leaders at the factories in 1890 were probably equally impressed by the dynamo. Then almost nothing happened for twenty years. The lineshaft holds things in place.
The lineshaft this time is in some ways worse, and David saw it coming. In his last paragraph he points out that a firm’s information structures are sunk costs like a factory, but unlike a factory they don’t physically depreciate, so you can’t count on time to create the occasion to redesign them. In 1890, it was physical infrastructure made with steel that depreciates - so each passing year made the math a little bit better. But not so this time - the edge case knowledge stored in the heads of individual employees is not depreciating, but as more gets automated around them, it becomes even more valuable as the linchpin holding together the whole system.
Which is counterintuitive from an entrepreneur’s perspective - the work that is most valuable to automate, that can for the first time be automated, is the complex, judgment-laden work that involves lots of steps. The simple stuff does not have as high a ROI.
You can watch this play out in shared services. Those teams have been outsourced for twenty years, documented for SOX, and metered to the invoice. Everything about them says easy. But everything specifiable left in 2011, and what’s still there is the exception pile, which is exactly the stuff nobody can hand you. Every previous wave selected for legibility. What’s left is illegible and in the heads of named human beings who would rather stay employed than not.
Nobody rebuilds a working factory
It rarely is ever worth replacing the working factory. The upfront investments are usually too big, the risk of it failing is too great, and the gains get a big discount rate because of time and how far out they are. People replace load-bearing systems only occasionally, and mostly under duress: either the old system is reaching end of support, there’s a merger, a regulator forces their hand, there’s a big outage, or the people who were doing the job leave the workforce, which makes retirement the most reliable digital transformation strategy.
And they definitely don’t do it in a downturn. While downturns provide the opportunity to make change, and give folks political cover to make unpopular decisions, they also raise the cost of capital significantly, and lower the risk tolerance. And so the math doesn’t math.
New factories
The obvious rejoinder is that nobody has to rebuild the factory. Someone else builds a new one instead, and that’s how this resolves.
As an entrepreneur and angel investor, I know that story very well. The story of the place I live - Silicon Valley - is exactly this, and it’s also the conversation of every coffee shop in my neighborhood. So, I’m the first person to believe in this narrative. I do think that’s how the change will happen.
But look at the Fortune 100. For a lot of them, there’s no insurgent coming at all. The margins aren’t fat enough to make a new entrant rich, or the moat is a brand built over a century, or the business is dull enough that nobody’s trying to reinvent it. The competitive pressure that’s supposed to force the rebuild just isn’t there.
It’ll take a while
Diffusion of technologies always takes time, and this technology is not showing any signs that it’ll go that much faster. The technology itself is getting better at mind-blowing speed, but the people and systems are moving at the same pace as always. The revolutionary electric dynamo took four decades, and even the computer took many decades. It’s possible this wave moves faster because the tech is that much better, but I’m skeptical.
And, this may not be a bad thing. I’m not against the person doing the 11 step process with the undocumented workaround. It’s inefficient, sure. It’s also fine. She’s good at something nobody ever wrote down, and the world will survive her doing it for another decade. Society has trouble with rapid change, and the time this buys is probably worth more than the productivity it costs.
So the SVP will keep asking. The demos will keep getting more impressive. And the ops team will keep doing the work by hand, next month and the month after that. Not because the work is hard, but because the only people who know how are the ones it would replace.

