INDUSTRIAL EDGE COMPUTING SERIES

Combining agentic AI with distributed orchestration at the industrial edge

Agents that move workloads, recover from faults and patch nodes at the industrial edge are technically possible. This paper sets out what a buyer has to verify before letting them act, starting with whether the plant's security and OT teams will allow it without approval.

Harinderpal Hanspal · LinkedIn · hans@thing.company · About 10 min read · 5 sections · References

Request a Sprint Assessment A thirty-minute call about your own decision, not a purchase.

Executive summary

Vendors now describe edge infrastructure that runs itself. An agent watches the nodes, moves workloads when demand shifts, routes around a failed box, and patches or isolates a node when it sees a threat. Distributed orchestration does the work underneath: it finds devices, balances load, and deploys containerized services across many sites.

Most of that is technically possible.

The question a buyer has to answer is a different one. Which of these actions will the plant's operations, security and OT teams let the agent take without a person approving it first? The answer sets what the product is worth to that buyer, and no architecture diagram contains it.

This paper describes the combination in plain terms. It then takes the main claims made for it one at a time: what the technology does, what is established, what is still open, and what a buyer should test before signing. Security is the sharpest case. US federal guidance for operational technology (OT) tells operators to test patches under field conditions before installing them, on a test system where possible. Real-time autonomous patching works against that guidance, and against the way most plant security teams operate.

The last section covers the commercial questions the architecture cannot answer, because they belong to people who never see it.

What the combination is

Agentic AI

A rule-based system does what its rules say. Anything outside the rules waits for a person. An agentic system is given a goal and chooses its own actions toward it: it observes the state of the edge estate, decides, acts, and checks the result.

That is a description of the category. Whether a given product behaves this way at a given site is an empirical question, and the rest of this paper treats it as one.

At the edge, the actions on offer are concrete:

  • Allocating compute, storage and network capacity across edge nodes as demand moves.
  • Detecting a fault, rerouting tasks, and restarting or replacing a failed service.
  • Setting the priority of traffic between edge nodes and central systems.
  • Responding to a suspected intrusion by patching or isolating a node.

Distributed orchestration

The agent decides. Orchestration carries the decision out across devices that are widely distributed and carry uneven workloads. In an industrial edge environment it covers three jobs:

  • Device discovery and management: detecting new devices such as sensors and robots, registering them on the network, and bringing them into the operational workflow.
  • Task offloading and load balancing: spreading computational tasks across devices according to the resources each has free, so no single node is overloaded.
  • Service deployment: rolling out services such as real-time analytics and predictive maintenance across edge nodes as containerized applications, which package a service with what it needs to run, so it can move between nodes more easily than software installed directly on each machine.
Agentic AI and distributed orchestration at the industrial edge Layer stack showing agentic AI above distributed orchestration, which deploys containerized services onto edge nodes and devices, with the AI and orchestration layers marked as the combination the paper argues for. L4 INTELLIGENCE Agentic AI distributes workloads, self-heals, responds to threats L3 COORDINATION Distributed orchestration discovers devices, balances load, deploys services L2 Containerized services real-time analytics, predictive maintenance L1 Edge nodes and devices compute, storage, network; sensors and robots The combination capabilities neither layer delivers alone
Agentic AI and distributed orchestration at the industrial edge

The split matters to a buyer. Orchestration is deterministic: given the same instruction, it does the same thing, and it can be tested like any other infrastructure software.

The agent is the part that chooses. Most of the risk sits in that choice, and so does most of the claimed value.

Testing this against your own decision? Request a Sprint Assessment

Why buyers look at it

The number behind the interest is downtime. Siemens estimates that unplanned downtime costs the world's 500 largest companies about 11 percent of their revenue, roughly $1.4 trillion a year, and that an hour of downtime at a large automotive plant costs $2.3 million [1]. The figures are Siemens' estimates, extrapolated from 181 online interviews with maintenance, engineering and IT professionals between April 2019 and March 2023, and Siemens sells predictive maintenance. Under our evidence standard they rate Benchmarked: secondary evidence, useful for direction and scale.

The same number cuts the other way. An agent that isolates the wrong node, or restarts a service in the middle of a shift, causes the downtime it was bought to prevent, at the same cost per hour. A buyer who uses the downtime figure to justify the purchase has to apply it to the agent's mistakes as well.

Most of what follows comes from that symmetry.

Five claims, and what a buyer should verify

The companion paper on edge infrastructure names eight barriers to industrial edge deployments. The combination described here claims to address four of them: fragmented systems, scaling, cybersecurity, and legacy equipment. They are grouped below by the action the agent takes, because an action is what a buyer approves or refuses.

Resource allocation across nodes

The claim. Industrial operations that depend on real-time data stall when capacity runs short in the wrong place. With orchestration underneath, an agent can give critical jobs what they need and spread routine work across whatever nodes are free. In a plant with changing workloads, predictive maintenance models might get priority while batch reporting moves to idle nodes. The same mechanism is offered as the answer to fragmentation and scale: new devices and nodes join, and the system absorbs them without interrupting running work.

What is established. Moving a containerized analytics job between nodes rarely touches the physical process, so this is the lowest-risk action in the set. The earlier version of this paper also said that deployment times drop. No before-and-after measurement supports that, and the claim has been removed.

What to test. Measure two numbers on the buyer's own estate: the time from a device's arrival to its first running workload, and node utilization under peak load. Agree the baseline before the pilot starts, or any result will look like an improvement.

Self-healing at remote sites

The claim. Remote or disconnected sites, such as offshore energy platforms or remote logistics hubs, are hard to fix quickly. In a logistics network, an edge node that tracks shipments fails with a hardware fault. The agent detects the failure and reroutes data to other nodes, so monitoring continues without anyone traveling to the site.

What is established. Rerouting a data flow around a failed node is standard practice in distributed systems. What self-healing saves depends on how long a failure takes to fix today, and we have not found a sourced mean time to repair for remote industrial sites. That is the number that sizes the benefit.

What to test. Take the repair times from the buyer's own maintenance records. Then ask the vendor for a written list of the failures the agent will recover from on its own, and the ones where it must call a person.

Traffic priority on the network

The claim. Decisions that depend on real-time data need low-latency links between edge nodes and central systems. In a smart grid, sensor data keeps supply and demand in balance. An agent can give that data priority and move less urgent traffic before congestion builds.

What is established. Reprioritizing traffic classes is routine network engineering. Some latency budgets in this domain are hard limits, set by standards: the turnkey platform paper cites the roughly three millisecond budget IEC 61850-5 sets for protection trip messages in substation automation, which applies to its higher performance classes (P2 and P3); the lowest class (P1) allows 10 milliseconds. An agent that makes decisions should not sit in the path of traffic with a budget like that.

What to test. Confirm in writing which traffic classes the agent may reprioritize, and that protection and safety traffic are outside its reach.

Patching and isolation on OT assets

The claim. The earlier version of this paper said that an agent watching for threats and deploying patches or isolation in real time "cuts the risk of a breach," and that edge systems therefore "stay secure and compliant." That is two claims, and they need to be separated.

What is established. The technical half holds. An agent can detect an anomaly and trigger a patch rollout or a network isolation through the orchestration layer.

The operational half is where the purchase is decided. NIST's guide to OT security tells operators to deploy patches quickly, but only after testing them under field conditions, on a test system if possible, before installing them on the OT system [2]. The same guide notes that availability concerns may push patching until the next planned operational downtime, and that operators without the means to test patches themselves depend on the vendor to validate them. Where deployment would require an unacceptable operations shutdown, it says compensating controls should be implemented instead. An agent that patches OT assets in real time skips the test step that guidance is built around.

Isolation carries the same risk. A node that feeds a controller is part of the process. Cutting it off can stop the line it was meant to protect.

We found no evidence that autonomous patching reduces breach incidence on OT assets, so the claim has been removed. The compliance claim has been removed as well. Compliance is decided by the regulation and the auditor, and a technical action does not produce it automatically. The earlier version illustrated this section with a healthcare facility; that example is out of scope for an industrial paper and has been dropped.

What to test. Ask the plant's security lead and OT lead, separately, which of three modes each would approve: the agent recommends and a person acts; the agent acts on IT-side edge nodes only; the agent acts on OT assets. Ask per asset class, because the answer can differ between a historian server and a line controller. If the two leads give different answers, the purchase will stall at the review between them.

Legacy equipment and protocol conversion

The claim. Older devices do not have to be ripped out. The earlier version said the agent bridges legacy and modern systems and handles protocol conversion on its own, so a site can modernize gradually.

What is established. Conversion between legacy fieldbus protocols and modern systems is done today by deterministic gateways and adapters. The newer pattern places an agent above those adapters: the adapter translates, and the agent decides what to read or write. Research on this pattern exists. One 2026 study built adapters for Modbus, MQTT/Sparkplug B and OPC UA and ran them through 870 test runs, against simulated endpoints rather than plant equipment [3]. We found no published evaluation of an agent handling protocol conversion on a production brownfield line. The claim is open.

What to test. Run the integration against the buyer's actual controllers in a test cell. Get a list of every write the agent can issue through each adapter, and who approved each one.

The decision the architecture does not make

All five claims can hold up and the purchase can still stall. The questions that decide it are commercial, and the people who ask them do not read architecture diagrams. For agentic systems at the edge, five come up. Our page on validating agentic and physical AI sets out how each is scored.

  • Autonomy tolerance. Which actions the buyer will let the agent take without approval, per asset class. A buyer who will accept recommendations only will not pay for a product priced on autonomous action.
  • Governance readiness. Whether the buyer can set, see and change the agent's permissions, and whether every action is logged against the agent that took it. Governing agents in production covers the mechanisms.
  • A falsifiable success criterion. A number agreed before the pilot that could come out wrong, such as recovery time for a named class of failure at a named site. "Improved resilience" cannot fail, so it cannot prove anything.
  • Budget and ROI ownership. Edge infrastructure can sit between the IT and OT budgets. One named person has to own the line item and the downtime saving claimed for it.
  • Liability and safety. Once an agent can isolate a node that feeds a controller, the buyer's safety function and legal counsel join the purchase. Who pays for a wrong action belongs in the contract. Validating physical AI covers the liability model in full.

Each of these is answered by finding the person who holds it and asking them. Testing the software answers none of them.

References

  1. Siemens (Senseye Predictive Maintenance), The True Cost of Downtime 2024. The figures are Siemens' estimates, based on 181 completed online interviews, April 2019 to March 2023, with maintenance, engineering and IT professionals at large industrial organizations. Vendor-published, and rated Benchmarked for that reason.
  2. Keith Stouffer et al., Guide to Operational Technology (OT) Security, NIST Special Publication 800-82 Revision 3, National Institute of Standards and Technology, September 2023. The patching passages are in section 6.2.11 and in the OT discussion of control SI-2 (flaw remediation).
  3. Melwin Xavier, Melveena Jolly, Vaisakh M A and Midhun Xavier, IndustriConnect: MCP Adapters and Mock-First Evaluation for AI-Assisted Industrial Operations, arXiv:2603.24703, March 2026. A preprint (submitted 25 March 2026), evaluated against mock endpoints in a mock-first workflow.

What is not yet sourced

  • Deployment time. No baseline exists for how long an edge deployment takes today, so no reduction is claimed.
  • Repair time at remote sites. No sourced mean time to repair was found. The buyer's own records are the better source in any case.
  • Autonomous patching and breach incidence. No evidence was found that autonomous patching lowers breaches on OT assets. The claim was removed, not weakened.
  • Autonomous protocol conversion in production. Open. The only evaluation found ran against simulated endpoints.
  • The five commercial questions are Thing Company methodology, not an external benchmark.

About Thing Company

Thing Company is an independent market validation practice for industrial technology. We test whether a buyer exists at a price that works.

Agentic systems raise a commercial question next to the architectural one: how much autonomy a buyer will grant, and what they will pay for it. Buyers answer that. The design does not. How we work sets out the evidence standard, and the Sprint is the instrument that applies it.

Working on an agentic or edge bet? Start with a Sprint Assessment.