A year ago, “agentic AI” felt like the next logical step in the evolution of artificial intelligence.
Today, it’s everywhere.
Large enterprises are deploying hundreds of AI agents across operations. Vendors are promising autonomous workflows, self-solving systems, and digital co-workers that can take on everything from customer service to field support.
On paper, it sounds like the breakthrough industry has been waiting for. In practice, it’s more complicated.
“An AI agent, in effect, is an autonomous digital co-worker that is capable of performing tasks and actually accomplishing things similar to how a human would work in a seat today,” says Cooper Schorr, senior director of North American sales at Octonomy AI, an agentic AI startup based in Cologne, Germany.
That’s the promise. But in industrial environments—where accuracy isn’t optional—that promise is running into hard limits.
The gap between promise and performance
For manufacturers, OEMs, and service organizations, the appeal of agentic AI is straightforward. There’s a significant amount of work surrounding complex equipment that doesn’t involve physically fixing anything. Things like opening and managing tickets, documenting field service work, or identifying parts are the kind of tasks companies want AI to handle—freeing up skilled technicians to focus on what they do best.
At the same time, OEMs are looking to automate lower-level service interactions entirely.
“If a question is as simple as, ‘What part do I need for X, Y, and Z?’ that’s knowledge work that today requires someone to go into Salesforce or ServiceNow and reply manually,” he says. “That’s agentic work that can be done a lot faster by artificial intelligence.”
But as many companies are discovering, getting AI to actually deliver on that promise is another matter.
The core issue, according to Schorr, is structural.
“Most AI systems are built to generalize across domains… and in industrial work, the opposite needs to be true.”
Every piece of equipment has its own specifications, fault codes and operating logic. Generalization—the strength of modern AI—is exactly what creates problems. And the consequences are measurable.
“These tools typically cap at about 40 to 60 percent accuracy at best on complex technical content,” Schorr says. “That’s fine for drafting an email. It’s not fine for dispatching the right part to a field site or walking a technician through a repair on a six-figure asset.”
At that level of performance, the issue transitions from inefficiency to risk.
The architectural problem
The deeper failure point isn’t just data quality or training; it’s how AI systems process information.
Most platforms follow an approach similar to this: They convert documents into text, analyze it and then apply reason. It’s a system that works well for written content but it breaks down quickly in industrial contexts.
“The moment you do that on a schematic or a wiring diagram, you lose the spatial relationships that make the diagram mean anything,” Schorr says. “A label without a line connecting it to a diagram is just a word.”
In other words, much of the knowledge that matters in industrial environments isn’t purely textual. It’s visual, spatial, and contextual. And once that structure is stripped away, the AI is no longer working with the correct information.
Confidence without correctness
These types of accuracy challenges are compounded by another, less obvious issue: how AI systems handle uncertainty. Hallucination—AI generating incorrect or fabricated answers—is a widely discussed issue. But Schorr argues the problem runs deeper.
“Hallucination in general-purpose AI is not a bug that the industry is racing to fix. It’s a design choice,” he contends. Why? Because the models are trained to behave in ways users prefer.
“These models are trained on human feedback, and humans prefer confident, agreeable-sounding answers over cautious ones,” he says. “The models learn to sound sure even when they’re wrong.”
In low-stakes applications that’s manageable, but in an industrial setting, it’s a liability.
“Wrong in an email is a problem, but often not a catastrophic one,” Schorr notes. “Wrong when you’re trying to figure out how to fix a $600,000 piece of equipment is a much bigger deal.”
The real risk isn’t just that the system is wrong—it’s that users don’t know when it’s making a mistake.
Despite these limitations, many organizations continue to invest heavily in AI pilots, often with underwhelming results.
“What ends up happening is that in 18 months, they realize they’ve solved almost nothing,” Schorr says. “They’ve really just run a pilot for show.”
He says part of the problem is how these systems are evaluated.
“Vendors are pushing people to run demos on the cleanest data and the simplest use case and calling it a win,” he says. But that creates a misleading picture of performance—one that doesn’t hold up in real-world conditions.
A different way to evaluate AI
Instead of testing ideal scenarios, Schorr recommends starting at the opposite end. That means start with the failure mode, not the feature set. He says if you really want to see what AI can do for you, test the hardest use case that has the most complex documentation to push the system to its limits.
“Pick the gnarliest technical manual, the densest schematic, the legacy document that’s been sitting on SharePoint for 15 years and then evaluate how the AI behaves when it fails” he says.
Is it willing to say, ‘I don’t know’? If it’s not, you might find yourself in in real trouble down the road. The goal isn’t to prove that AI works—it’s to understand where it breaks.
A shift toward specialization
As these limitations become clearer, a divide is emerging in the AI landscape. On one side are general-purpose models designed for scale, on the other are specialized systems built for specific problem classes. They key, according to Schorr, is depth over breadth—a specialist can pick one problem and go deeper than anyone else can economically justify.
As an example, Schoor point out that his company, Octonomy AI, has decided on taking this exact path. The problem is complex technical documentation, and instead of converting everything to text, the platform processes visual content directly.
“We built a different ingestion model, designed to process visual content the way a human visual cortex does,” Schorr explains, saying the result is an AI agent that interprets diagrams and schematics while maintaining context across large knowledge bases to deliver precise, situation-specific answers.
From information to action
That shift enables a different kind of output. Instead of pointing users to documents Octonomy’s agent pulls from manuals, schematics, and engineering drawings and delivers a short, specific answer tied to that moment, that machine—not simply a link to a 300-page PDF with a ‘good luck.’
And critically, the system captures what happens next.
“The outcome of every interaction—what worked, what part was actually needed—gets written back into the customer’s systems of record,” Schorr says, creating something many organizations have struggled to achieve: institutional knowledge that grows over time—and remains under their control even when the human experts exit the organization.
For many manufacturers, the most significant opportunity isn’t new data. It’s better use of existing information.
“Organizations spend millions of dollars and hundreds of thousands of hours building that documentation,” Schorr says. “But no one uses it because it’s so dense. Everyone just calls John or Jane who’s been there for 30 years and knows everything,” he says.
What buyers should be asking
As interest in AI continues to grow, Schorr emphasizes the need for more disciplined evaluation. By now, most organizations will ask if it has control over how the system behaves and if it can be trusted in high-stakes scenarios.
The more important question that rarely gets asked is: Who owns the knowledge it generates?
“If every fix and every answer lives in a vendor’s black box, the switching cost becomes permanent,” he says. “The knowledge base is held hostage.”
The next phase of AI adoption in manufacturing won’t be defined by broader deployment. It will be defined by those that manage to extract better reliability and accuracy from the AI models they deploy.
“Trust comes through reliability and predictability, not marketing,” Schorr says. “If a system can’t perform in real conditions, it won’t be used—and it shouldn’t be.