THING COMPANY PAPERS

Validating physical AI before the capital is irreversible

Physical AI fails the way software does, on an untested commercial hypothesis, and adds three questions software never had to score: the liability model, the sim-to-real gap at the buyer's site, and operational continuity.

Harinderpal Hanspal · LinkedIn · hans@thing.company · About 17 min read · 9 sections · Appendix · References

Request a Sprint Assessment A thirty-minute call about your own decision, not a purchase.

Executive summary

Physical AI now gets the kind of attention generative AI got first. At CES in January 2025, Nvidia's chief executive said the "ChatGPT moment" for general robotics was "just around the corner." In January 2026, Nvidia's press release quoted him saying "the ChatGPT moment for robotics is here," though on stage he said the ChatGPT moment for physical AI was "nearly here." The surveys moved with him: Deloitte's 2026 enterprise AI survey found 58% of companies reporting at least limited use of physical AI, a figure it expects to reach 80% in two years.

When enthusiasm runs that high, large commitments get made on thin evidence. This paper is for the people making them: an operations leader, usually a plant manager or VP of operations, deciding on a fleet, an original equipment manufacturer (OEM) weighing an autonomous product line, an investor deciding whether to fund one.

Physical AI fails the way software initiatives fail. The technology usually works; the untested part is whether the buyer who has to sign has the problem at the urgency and the price the business case assumed.

What physical AI adds is three questions software validation never had to score, and any one of them can stop a purchase. The first is the liability model: who answers for it when the system injures a worker, damages equipment, or stops a line, and whether the buyer's safety and legal teams accept that answer. The second is the gap between simulation and the buyer's own site, where adapting the system to each installation is either manageable or turns a product company into a services business. The third is operational continuity, meaning whether maintenance, recalibration, and failure recovery fit a production schedule that does not stop.

Autonomy level turns out to be a pricing decision as well. Industrial buyers adopt autonomy in steps, and a system priced for full autonomy stalls with a buyer prepared to approve only a recommendation.

A strong demonstration with the liability question still open is an unvalidated hypothesis.

The sections below describe the evidence that settles each dimension, and the appendix lists the questions we ask before a fleet commitment.

The inflection is real, and so was every previous one

The signals go beyond keynote language. Nvidia's chief financial officer told investors in February 2026 that physical AI had contributed more than USD 6 billion of revenue in the company's fiscal 2026, still under 3% of the total. A paid MarketsandMarkets report, as announced in its press release, sizes the physical AI market at USD 1.50 billion in 2026, rising to USD 15.24 billion by 2032, a compound annual growth rate of 47.2%. The two figures do not reconcile: one company's physical AI revenue already exceeds the report's estimate for the whole market in 2026, which shows how much the definition of "physical AI" drives the number. In manufacturing, a Manufacturing Leadership Council survey in early 2025, cited by Deloitte, found 9% of responding manufacturers using physical AI and 22% planning to within two years. None of those numbers tells you whether a specific buyer will sign, at a price that returns the investment, through a channel that reaches them. A trade-show demonstration cannot answer that either. Primary research with the people who have to say yes can.

In our experience, every major technology transition since the early 1990s attracted capital on the strength of technology that worked: client-server computing, the internet build-out, cloud, and the move to the edge. Each produced a similar split. A minority tested the commercial hypothesis before committing. Most committed first, building the sales motion and signing the channel before anyone checked, and learned later that the buyer, the price, or the channel was wrong.

The software side of the current wave already has that shape. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls. Read closely, each of those causes is a validation failure. Costs escalate when integration and autonomy assumptions were never tested. Value stays unclear when nobody confirmed the operational outcome with the buyer. Risk controls fall short when nobody scoped the trust threshold or the liability model. Physical AI carries the same exposure with more capital behind it, since its mistakes are made in hardware and site engineering and safety as well as in code.

The deployment gap in manufacturing is visible too. In Deloitte's 2025 smart manufacturing survey of 600 executives at US manufacturers with annual revenue of USD 500 million or more, 92% said smart manufacturing would be their main driver of competitiveness over the next three years. Only 29% were using AI or machine learning at the facility or network level. We read the distance between those two figures as commitments waiting on evidence more than as a shortfall in the technology.

Why interest is not a business case

Humanoid robots show the pattern clearly, because the demand behind them is real. Barclays counted 21 new humanoid models introduced in 2025. The International Federation of Robotics puts the operational stock of industrial robots at about 4.66 million units at the end of 2024. Barclays points to the same pressures: aging populations, and younger generations less inclined to take repetitive, physically demanding jobs. A plant leader with a shift they cannot fill is right to take the meeting.

A labor shortage proves urgency and little else. It says nothing about whether the buyer will pay a price that returns the investment, whether the system's uptime fits a continuous production schedule, or whether maintenance and recalibration will eat the labor savings.

Pilots hide that last question. A pilot site usually gets vendor engineers on call, a motivated internal champion, and attention that a fleet of dozens of sites will never receive, so the support load measured during a pilot is lower than production will need by design. The commercial question is the maintenance load left once the vendor's pilot-grade attention moves on, and who inside the buyer's organization absorbs it.

The sequence is predictable enough to describe in advance. The system is bought against a labor number, run against a production schedule nobody tested it on, and parked within a few quarters, because the shortage was measured and the economics were assumed.

Dimension one: the liability model

A wrong recommendation from software is a bad suggestion. A wrong physical action can injure a worker, damage equipment, or halt a line, and that changes who gets a say in the purchase.

For analytics software, the buyer is usually an operations or IT leader. For an autonomous physical system, the buyer's environment, health, and safety (EHS) function (HSE in energy and oil and gas) and its legal counsel enter the decision, and either one can stop it. Most vendor pitches were never built to answer what they evaluate.

Under our evidence standard, "we are working through the liability model" rates as an Assumption, since it is one stated as progress. While the question stays open, a good demonstration will not move the deal. It stalls in a review the vendor did not know existed.

A liability structure the buyer has accepted

A confirmed liability structure is one the buyer's own legal and EHS teams have reviewed and accepted. It answers four questions:

  • Who carries insurance, and for which failure modes.
  • What the failure-mode analysis covers.
  • What human oversight the deployment requires.
  • Where responsibility sits when the system operates unsupervised.

The test is simple to run. Ask the buyer to name the person who would sign off on the liability terms, then interview that person. If nobody can be named, the hypothesis has a dimension graded Assumption under a demonstration graded Verified.

Settle this before the go-to-market scales. Whether a pilot converts or gets admired and then declined is often decided here, by people the vendor never presented to.

Where liability sits in the architecture

Liability follows actuation, and actuation happens at the protocol layer where software meets brownfield equipment. That gateway is the last point where every action is still a request that can be classified by consequence, gated, logged, and refused.

CISA and eight partner agencies advise (guidance of 3 December 2025) that where AI actively updates control logic, systems should add human-in-the-loop intervention points, build failsafe mechanisms, and include how to bypass or replace the AI system in functional safety and incident response processes. They also write that humans are responsible for functional safety. It is guidance, not regulation. Four questions follow for any vendor: is each action classified by consequence before it runs, does every route to the device pass that check, is the list of what an agent may do built at the moment it asks, and what lowers an agent's autonomy after repeated failures? A capability list kept as documentation cannot be trusted to match what the runtime will dispatch. Our paper on governing agents in production sets out those questions and the controls around them.

Where liability sits: the protocol gateway Architecture showing a request from an agent or a person passing through a protocol gateway that checks the capability catalog, classifies the action by consequence against the configured autonomy level, logs the decision, and permits or refuses it before the protocol bridge reads or writes brownfield equipment, with a circuit breaker lowering autonomy after repeated failures. PROTOCOL GATEWAY REQUEST REGISTERED LIMITS LOGGED LOWERS PERMITTED READ OR WRITE FAILURES Agent or person requests an action Capability catalog built at dispatch Classify action by consequence Autonomy level configured Action log every decision Gate permit or refuse Circuit breaker on repeated failure Protocol bridge read sensor, write setpoint Brownfield equipment setpoints on machines LEGEND requester, step, store, equipment liability boundary failure signal
Where liability sits: the protocol gateway

A buyer's legal team will eventually ask what touched the machine, under whose authority, and with what checks. A vendor who treats the protocol layer as an integration detail has no answer to give.

Testing this against your own decision? Request a Sprint Assessment

Dimension two: the sim-to-real gap at the buyer's site

Simulation and lab results describe performance in a controlled environment. The commercially relevant number is performance at a specific installation, with its own lighting, surfaces, and dust, equipment of mixed vintage, its layout, and the undocumented workarounds the night shift relies on.

Extra lab work cannot settle that, because the gap belongs to the site as much as to the model. It can only be measured by running the system in the buyer's environment, which makes it a validation question with a commercial answer.

When every site needs its own engineering

If closing the gap requires bespoke engineering at each new site, the company is a services business with product-company pricing. Margins shrink with each deployment and timelines stretch, until the scalability story behind the business case is no longer true.

The failure rarely shows in any single deal, because each one closes with "some integration work." It becomes visible in aggregate, when deployments have consumed the whole engineering roadmap.

The exposure grows with adoption: every manufacturer in the Manufacturing Leadership Council's 22% will find its own per-site gap, during validation or during the contract.

Measuring the gap where the buyer operates

A findings brief on a physical AI company should state the measured or bounded adaptation cost at representative target sites, who bears that cost under the proposed contract, and what it does to unit economics at the quoted price. "We are working through site variability" is an Assumption. A per-site adaptation cost confirmed in a representative environment and priced into the model is Verified.

Where the gap was measured makes a difference. A result validated at the vendor's own demonstration facility says little about the buyer's plant, and the buyer's operations lead will usually ask about their own site as soon as the pitch moves past the simulation footage.

Dimension three: operational continuity

The third question belongs to the operations leader. Is the system's profile for maintenance, recalibration, and failure recovery compatible with a continuous production schedule?

A plant running three shifts cannot pause a line for recalibration when it suits the vendor. A system that needs a specialist on site to recover from a fault adds a staffing cost the business case may not include, and one whose failure stops the line, where the plant expected a fallback to manual operation, changes the risk the buyer is taking on.

These questions get answered with the operations leader who owns the line and the maintenance and reliability manager whose team will support the system, against the buyer's own shift pattern. An uptime figure from a pilot does not answer them.

What operations needs to see

  • A maintenance and recalibration schedule confirmed against the buyer's production calendar.
  • A failure-recovery path the operations team has walked through, including who responds and how quickly.
  • A support model that works without vendor engineers on call, with the internal owner named.

Pilot success changes who approves

A physical AI pilot impresses the people closest to it. Production brings in approvers who never saw it: safety and legal once the system runs without the vendor watching, multi-site operations once results have to carry across plants, and finance once the spend becomes a capital request. Identify those approvers, and the evidence each will require, while the pilot is still running. Our paper on the industrial buying committee covers how to map them.

Autonomy level is a commercial variable

Two vendors with the same use case can see very different sales cycles, because one asks the buyer to approve a system that recommends and the other a system that acts. The second request carries more risk and pulls in more of the buying committee, and it often comes out of a different budget.

The difference is who absorbs a bad call. When a system recommends, a person does. When it acts on a production schedule, a maintenance plan, or a supplier order, the buyer is trusting it with outcomes they are personally accountable for: throughput, safety, uptime. That trust does not grow with model accuracy. A buyer grants it at a specific level of autonomy, for a specific use case, at a specific site.

Industrial buyers adopt autonomy in steps. Gartner's Manufacturing Predicts 2026 report, dated 10 December 2025 and known here through published summaries, expects semiautonomous AI agents to orchestrate 10% of key production, quality, and maintenance use cases by 2030, up from 2%, with humans keeping final approval. Inside a single plant, the same manager may trust a system to flag an anomaly, hesitate to let it reschedule maintenance, and refuse to let it stop a line. Trust varies with the use case, the site, and the consequence of a wrong action, and most of all with how much autonomy is being requested.

We score trust as its own dimension in discovery interviews for that reason, on a one-to-five scale, with a strong signal at an average of 4.0 or above, the same threshold every other dimension uses. A low score means the product configuration has to change; pushing harder in the sales process will not raise it. The viable product at that account may be a version in which the system recommends and an operator confirms, with more autonomy earned later against an operating record.

The validation question becomes the level of autonomy at which this buyer will approve the system, and whether the commercial model still works at that level. Confirm it before the product is built and priced for full autonomy, because a vendor who packages ahead of the buyer's step creates its own stall. Price and package for the step the buyer is on.

What doubling down should require

The scale decision is the hardest moment to bring evidence to. By then the team that built the program has the most conviction and the least incentive to press on the three questions above. Capital is deployed, people are hired, and leadership has said yes in public.

The reluctance is structural. Nobody inside a committed program can challenge its direction without paying an organizational cost, and that cost shapes what people are willing to say out loud. Investors and strategic partners, meanwhile, will ask for exactly the evidence the team is least positioned to produce. They have read the return data on AI pilots. One widely read example is MIT's Project NANDA, in its preliminary July 2025 report on generative AI in business, which found that about 5% of integrated AI pilots were extracting millions in value. It rests on 52 interviews, a survey of 153 leaders and a review of more than 300 public initiatives, and its authors call the figures directionally accurate. It concerns generative AI, not physical AI, so it shows the scrutiny a scale decision meets and does not size the odds for a physical AI program. The distance between internal confidence and that scrutiny is the evidence that is missing.

A scale decision should rest on an independent assessment with three parts:

  1. Primary evidence from the economic buyers and operations leaders at representative target accounts, including the plant controller who checks the value case. The champion who hosted the pilot is not enough.
  2. The three physical AI dimensions scored alongside the standard commercial ones: problem confirmation, willingness to pay, market size, access to the decision maker, the alternatives the buyer compares the system to, and solution fit.
  3. A named verdict of Proceed, Pivot, Reset, or Stop, with the evidence behind every rating written down before the decision meeting.
What a physical AI verdict has to score A physical AI assessment scores the six standard commercial dimensions alongside three dimensions software validation never scored, the liability model, the sim-to-real gap at the buyer's site, and operational continuity, each able to block a purchase on its own, before a named verdict of Proceed, Pivot, Reset, or Stop. SIX COMMERCIAL DIMENSIONS PHYSICAL AI ADDS THREE 01 Problem confirmation 02 Willingness to pay 03 Market size 04 Access to the decision maker 05 Alternatives the buyer compares 06 Solution fit each one can block a purchase alone The liability model safety and legal teams can veto the deal Sim-to-real gap at the site adapting to each installation must pay Operational continuity fits a continuous production schedule Named verdict Proceed, Pivot, Reset, or Stop LEGEND standard dimension physical AI dimension verdict
What a physical AI verdict has to score

If the evidence supports scale, the fleet conversation and the funding decision both get stronger. If it does not, the finding arrives at validation prices, and in physical AI, deployment prices include hardware, site engineering, and safety exposure.

The evidence standard, briefly

Every claim in that assessment carries a rating. Verified means confirmed by scored interviews with the economic buyer at a specific price and use case. Benchmarked means supported by benchmarks, analyst data, or a technical champion who cannot sign a purchase order. An Assumption is a claim without that support, stated openly and given a validation path.

Most claims teams call validated turn out to be Benchmarked, and by that definition so is every market figure in this paper: the surveys and forecasts at the top show the category is real without saying whether your buyer will pay your price. Assumptions are allowed. What the standard requires is knowing which claims are Assumptions before committing capital to them.

Appendix: physical AI validation questions

Liability and safety

  1. Who at the buyer would sign off on the liability terms, and has that person been interviewed?
  2. Have the buyer's legal and EHS teams reviewed the liability structure, and what did they accept?
  3. Who carries insurance, and for which failure modes?
  4. What human oversight does the deployment require, and where does responsibility sit when the system runs unsupervised?
  5. Are actions classified by consequence at the protocol layer, with reads and writes on different terms?
  6. Does anything lower autonomy automatically after repeated failures?

Sim-to-real at the site

  1. Has adaptation cost been measured or bounded at representative target sites, or only at the vendor's facility?
  2. Who bears that cost under the proposed contract?
  3. What does per-site adaptation do to unit economics at the quoted price?
  4. How much of the engineering roadmap do current deployments already consume?

Operational continuity

  1. Is the maintenance and recalibration schedule confirmed against the buyer's production calendar?
  2. When the system fails, does the line stop or fall back to manual operation?
  3. Who responds to a fault, and how quickly, once vendor engineers are no longer on call?
  4. Who inside the buyer's organization owns the support load, by name?

Autonomy and trust

  1. At what level of autonomy will this buyer approve the system?
  2. What is the average trust score across discovery interviews, and does it reach 4.0?
  3. Does the commercial model work at the autonomy level the buyer will approve?

Commercial fundamentals

  1. Has the economic buyer confirmed the problem, or only the engineer who liked the demonstration?
  2. Does the labor or productivity case survive the maintenance load the system creates?
  3. Which budget pays for the system, who owns that budget, and does the pricing model fit how it approves spending?
  4. Beyond the people who ran the pilot, who will approve production, and what evidence does each of them require?

References

  1. NVIDIA, "CES 2025: AI Advancing at 'Incredible Pace,' NVIDIA CEO Says," NVIDIA Blog, January 2025. https://blogs.nvidia.com/blog/ces-2025-jensen-huang/
  2. NVIDIA, "NVIDIA Releases New Physical AI Models as Global Partners Unveil Next-Generation Robots," press release, 5 January 2026, quoting its chief executive: "The ChatGPT moment for robotics is here." https://nvidianews.nvidia.com/news/nvidia-releases-new-physical-ai-models-as-global-partners-unveil-next-generation-robots Fortune, 6 January 2026, reports that in the keynote he said the ChatGPT moment for physical AI is "nearly here." https://fortune.com/2026/01/06/nvidia-jensen-huang-chatgpt-moment-for-robotics/
  3. Deloitte, "The State of AI in the Enterprise, 2026," survey of 3,235 leaders in 24 countries conducted August to September 2025: "More than half of companies (58%) report at least limited use of physical AI today, and that figure is set to reach 80% in two years." https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html
  4. NVIDIA, fourth-quarter fiscal 2026 earnings call, 25 February 2026, transcript. Chief financial officer Colette Kress: "Physical AI is here, having already contributed north of $6 billion in NVIDIA revenue in fiscal year 2026." https://s201.q4cdn.com/141608511/files/doc_financials/2026/q4/NVDA-Q4-2026-Earnings-Call-25-February-2026-5_00-PM-ET.pdf Fiscal 2026 revenue of USD 215.9 billion is from NVIDIA's results release (Form 8-K). https://www.sec.gov/Archives/edgar/data/1045810/000104581026000019/q4fy26pr.htm The "under 3%" share is our calculation: 6 divided by 215.9 is 2.8%.
  5. MarketsandMarkets, "Physical AI Market worth $15.24 billion by 2032," press release via PR Newswire, 3 April 2026. https://www.prnewswire.com/news-releases/physical-ai-market-worth-15-24-billion-by-2032---exclusive-report-by-marketsandmarkets-302732794.html
  6. Deloitte, "2026 Manufacturing Industry Outlook," citing a Manufacturing Leadership Council survey of its respondents from early 2025. https://www.deloitte.com/us/en/insights/industry/manufacturing-industrial-products/manufacturing-industry-outlook.html
  7. Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
  8. Deloitte, "2025 Smart Manufacturing and Operations Survey," 600 executives at US manufacturers with annual revenue of USD 500 million or more and more than 1,000 employees, surveyed August to September 2024, press release via PR Newswire. https://www.prnewswire.com/news-releases/deloitte-survey-reveals-smart-manufacturing-is-driving-advantage-but-needs-focused-investment-and-implementation-302443462.html
  9. Barclays Investment Bank, "AI Gets Physical," Impact Series 14, 14 January 2026. https://www.ib.barclays/content/dam/barclaysmicrosites/ibpublic/documents/our-insights/impactseries14/Barclays Impact Series 14 - AI Gets Physical.pdf
  10. International Federation of Robotics, "World Robotics 2025 report: industrial robots," press release, 25 September 2025: "The total number of industrial robots in operational use worldwide was 4,664,000 units in 2024." https://ifr.org/ifr-press-releases/news/global-robot-demand-in-factories-doubles-over-10-years
  11. Gartner, "Manufacturing Predicts 2026: AI Agents, Digital Twins and the Race to Autonomous Operations," 10 December 2025 (paywalled; figure confirmed through published summaries, for example Synera's). https://www.synera.ai/analyst-study/gartner-manufacturing-trends-2026
  12. Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari, MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025," July 2025, preliminary findings from research between January and June 2025. The MIT-hosted PDF link now redirects to the project page (checked 5 October 2026); the copy cited is the one we read in full. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf Also reported by Virtualization Review, 19 August 2025. https://virtualizationreview.com/articles/2025/08/19/mit-report-finds-most-ai-business-investments-fail-reveals-genai-divide.aspx

What is not yet sourced

  • The pattern across technology transitions since the early 1990s is Thing Company's observation from its own work, not a measured study.
  • The NANDA figure (reference 12) comes from a preliminary generative AI study with a small, self-reported sample. It is context for the scrutiny a scale decision meets, not a physical AI benchmark.
  • The Gartner 2030 figure (reference 11) was confirmed through secondary summaries of the report, not the report itself.
  • Pilot support loads, EHS veto authority, and autonomy adopted in steps inside a plant rest on Thing Company's engagement experience, not on a published study.
  • The trust score threshold of 4.0 is Thing Company methodology, not an external benchmark.

About Thing Company

Thing Company is an independent market validation practice for industrial technology. We test whether a buyer exists at a price that works.

For physical AI, that means testing the liability model, the site, and the production schedule alongside the buyer and the price. How we work sets out the evidence standard, and the Sprint is the instrument that applies it.

Deciding on a fleet or a product line? Start with a Sprint Assessment.

Not ready to talk? Run the free five buyer questions check self-check from the toolkit.