PHYSICAL AI
Robots, machine vision and autonomous machines in plants, warehouses and energy sites fail the way software does, on a commercial hypothesis nobody tested. Physical AI adds three questions software never had to score, and any one of them can stop the purchase.
A working system is not the question. The untested part is whether the buyer who has to sign has the problem at the urgency and the price the business case assumed.
Pilots hide the rest. A pilot site gets vendor engineers on call, a motivated champion, and attention a network of dozens of sites will never receive. Production then brings in approvers who never saw the pilot: EHS and legal, the VP of operations over the other sites, and the plant controller once the spend becomes a capital request.
Each is scored on primary evidence from the people who can stop the purchase: the buyer's environment, health and safety (EHS) and legal teams, the plant manager who owns the line, and the maintenance and reliability team who will support the system.
Who answers when the system injures a worker, damages equipment, or stops a line? For an autonomous physical system, EHS and legal counsel enter the decision, and either one can stop it. A confirmed structure names who carries insurance and for which failure modes, what the failure-mode analysis covers, what human oversight the deployment requires, and where responsibility sits when the system runs unsupervised.
"We are working through the liability model" is an Assumption. A structure the buyer's legal and EHS teams have reviewed and accepted is Verified. The test: ask the buyer to name the person who would sign off on the liability terms, then interview that person.
Lab and simulation results describe a controlled environment. The number that matters is performance at one installation, with its lighting, surfaces and dust, equipment of mixed vintage, and the workarounds the night shift relies on. If closing the gap needs bespoke engineering at every site, the company is a services business with product-company pricing.
"We are working through site variability" is an Assumption. A per-site adaptation cost confirmed in a representative environment and priced into the model is Verified. A result from the vendor's own demonstration facility says little about the buyer's plant.
Do maintenance, recalibration, and failure recovery fit a production schedule that does not stop? A plant running three shifts cannot pause a line when it suits the vendor, and a system that stops the line on a fault, where the plant expected a fallback to manual operation, changes the risk the buyer is taking on.
A pilot uptime figure does not answer it. The evidence is a maintenance schedule confirmed against the buyer's production calendar, a failure-recovery path the plant manager and the maintenance and reliability team have walked through, and a support model that works without vendor engineers on call, with the internal owner named.
Industrial buyers adopt autonomy in steps. The same plant manager may trust a system to flag an anomaly, hesitate to let it reschedule maintenance, and refuse to let it stop a line. Trust is scored the same way as for agentic AI. Below the threshold, the viable product recommends and an operator confirms, and the commercial model has to work at that level. A system priced for full autonomy stalls with a buyer prepared to approve only a recommendation.
The revenue model is tested the same way. A robot or vision system can be sold outright from a capital budget (capex), run as a subscription or robot-as-a-service from an operating budget (opex), or both: the hardware bought once and the software and service paid every month. The evidence shows which one the buyer will approve, and who owns the service: the OEM, a dealer or an integrator.
Agentic AI in software raises the same question without the hardware: how we score agentic AI.
The buyer's legal and EHS teams have accepted the liability structure. The per-site adaptation cost is measured at a representative site and priced in. The maintenance schedule is confirmed against the production calendar, with the internal owner, or the regional integrator, named. All of it comes from scored interviews.
A strong demonstration, a pilot uptime figure, a result from the vendor's facility, and market forecasts are all real evidence. None of them confirms that this buyer's safety team, site, and shift pattern will accept the system.
Typical untested claims: the liability model is being worked through; site variability is being worked through; the pilot's support load will hold at every plant. A strong demonstration with the liability question still open is an unvalidated hypothesis.
The three physical AI dimensions are scored alongside the standard commercial ones, on evidence from the plant managers and VPs of operations who sign at representative target accounts or sites, not the champion who hosted the pilot. The Sprint ends in a recommended verdict of Proceed, Pivot, Reset, or Stop, with the evidence behind every rating written down before the decision meeting. For a deal, these three questions are the technical diligence items the thesis depends on.
The full argument, with the questions we ask before a network rollout, is in the paper: Validating physical AI before the capital is irreversible. The specimen brief shows what the scoring produces. For a first-hand note on where liability sits in a running system, see the protocol layer is where physical AI gets liable.
It is testing, with the people who can stop the purchase, whether robots, machine vision or autonomous machines will be approved and paid for at the buyer's site. It scores the standard commercial questions plus three that software never had to answer: the safety and liability model, the sim-to-real gap at the buyer's site, and operational continuity.
A working system is not the open question. The untested part is whether the buyer who has to sign has the problem at the urgency and price the business case assumed. A pilot site also gets vendor engineers on call, a motivated champion and attention a network of dozens of sites will never receive.
Production brings in approvers who never saw the pilot: environment, health and safety (EHS) and legal teams, the VP of operations over the other sites, the plant manager who owns the line, the maintenance and reliability team who will support it, and the plant controller once the spend becomes a capital request.
Either, or both. A robot or vision system can be sold outright from a capital budget, run as a subscription or robot-as-a-service from an operating budget, or split into hardware bought once and software and service paid monthly. The evidence shows which one the buyer will approve and who owns the service: the OEM, a dealer or an integrator.
One conversation to find out which of the three questions is still an assumption.
Or reach us directly at hello@thing.company