THING COMPANY PAPERS
Too successful to stop, too unproven to scale
Why industrial pilots succeed and never convert, why the people inside cannot make the call, and what a structured decision, including a defensible Stop, requires.
Harinderpal Hanspal · LinkedIn · hans@thing.company · About 21 min read · 10 sections · Appendix · References
Request a Sprint Assessment A thirty-minute call about your own decision, not a purchase.
Executive summary
Many industrial technology pilots succeed, and then nothing happens. The technology works, the results are encouraging, the champion is enthusiastic, and eighteen months later the initiative is on its third site, its second sponsor, and its fourth extension. Nobody has decided to scale it or to stop it.
We call that pilot purgatory: too successful to stop, too unproven to scale.
This paper is for the operators and executives holding a portfolio of pilots that will not convert, and for the vendors whose revenue depends on those pilots converting.
A pilot answers the least contested question in the initiative, whether the technology works under favorable conditions. The questions that decide conversion, about the other plants, the budget, and the frontline, were never in its scope. Purgatory then lasts because nobody convenes a decision. Stopping would require someone to say the commercial hypothesis was never validated, an admission with an organizational cost, while another extension feels like progress and costs nothing this quarter.
Meanwhile the cost of the wrong hypothesis spreads. Sales, support, finance, operations, legal, and channel each see their own symptom of the same upstream cause, and none of them is placed to name it. The people inside the initiative cannot make the call either, for reasons of incentive more than character, which is the gap an independent verdict fills.
The way out is a structured decision with criteria declared before the research, primary evidence from the economic buyer (usually the plant manager or VP of operations) and the machine operators on the frontline, and a verdict of Proceed, Pivot, Reset, or Stop. A defensible Stop is often the most valuable of the four. It protects the most capital, and it usually contains the sharpest brief for the next hypothesis.
The appendix is a diagnostic: the questions we ask about a pilot that has succeeded and not converted.
What a pilot proves, and what it never had in scope
A pilot is an experiment run under favorable conditions. It has one site, a motivated champion, a prepared environment, and a vendor with every incentive to make it work. When it succeeds, it confirms that the technology functions under those conditions.
That counts as evidence, on the least contested question in the whole initiative.
The questions that decide whether a pilot converts to production were never in its scope:
- Does the operational problem exist at the same urgency across the network, or was it acute only at the pilot site?
- Does the budget that approved a limited evaluation have the authority to approve deployment across the network? Usually it does not. An innovation budget has a light process and a champion with discretion, while a network capital request has a different owner and a higher standard of evidence.
- Will frontline workers at sites without a champion adopt the system?
- Does the solution generalize across sites with different equipment ages and workforce profiles?
- What support load remains once the vendor's pilot-grade attention moves on?
None of these are technology questions, and every one of them is still an assumption on the day the pilot ends.
Pilots leave them unanswered by design. A pilot is built to demonstrate feasibility, and network commitments are routinely made before any of the scale questions is answered. A pilot also brings in approvers it was never tested against: safety and legal, multi-site operations, and finance join only when production is on the table. Our paper on the industrial buying committee covers those seats.
The numbers describe a purgatory
The industry data shows a wide gap between intent and deployment.
- In Deloitte's 2025 smart manufacturing survey of 600 executives at companies with annual revenue of US$500 million or more and US headquarters or operations, 92% of manufacturers surveyed said smart manufacturing would be the main driver of competitiveness over the next three years. Only 29% were using AI or machine learning at the facility or network level.
- In Redwood Software's 2026 manufacturing outlook, a global survey of 300 manufacturing professionals, 98% of manufacturers were exploring or considering AI-driven automation, and only 20% said they felt fully prepared to use it at scale.
- MIT's Project NANDA, in its preliminary July 2025 report on generative AI in business, found that about 5% of integrated AI pilots were extracting millions in value, while the large majority showed no measurable impact on profit and loss. The authors call the figures directionally accurate: they draw on 52 interviews, a survey of 153 leaders and a review of more than 300 public initiatives.
The space between those numbers is full of pilots that succeeded and initiatives that did not convert.
Some organizations are now making the decision purgatory avoids. In S&P Global Market Intelligence's 2025 survey of more than 1,000 respondents in North America and Europe, the share of companies abandoning most of their AI initiatives rose to 42%, up from 17% a year earlier. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls.
Newer categories are entering the same pipeline with the same structure. Agentic AI and physical AI initiatives arrive with a technically impressive pilot, an enthusiastic champion, and an untested commercial hypothesis underneath. The category changes with each cycle, and the way initiatives stall stays put.
The pattern is older than AI
CB Insights' 2014 analysis of 101 startup post-mortems found that 42% cited no market need. Its March 2026 review of 431 venture-backed companies that shut down since 2023 found poor product-market fit as the primary reason behind 43% of the 385 whose failure reasons could be identified. The 2014 count includes every reason a post-mortem cited and the 2026 count assigns each company one primary reason, so the two do not form a precise time series. Even so, the stability of the finding is hard to ignore: tooling, talent, capital, and playbooks all improved over more than a decade, and the share of failures rooted in the market stayed where it was.
That persistence points to sequencing. The commercial hypothesis gets formed, funded, and acted on before anyone tests it, and no amount of downstream excellence fixes an untested upstream assumption. In our reading, few failed ventures lacked market need in the loose sense. What was missing was need at the specifics of one buyer, their urgency, the price, the channel, and the approval process. The loose version of a hypothesis survives every friendly conversation. The specific version can be proven wrong, and testing it is what separates market validation from market enthusiasm.
In our reading, cheaper tools do not change this. Lower model and build costs let more teams build before they test, so the untested hypothesis reaches the market faster. When runway is the constraint, test first the assumption whose failure ends the company, which is usually willingness to pay at the price the model needs. A Stop confirmed in month two leaves capital to act on. An unclear Proceed in month ten does not.
In our reading, each major technology transition has split along the same line. Client-server, the internet, cloud, and now edge and AI all drew investment into technology that worked. A minority confirmed buyer, price, and channel with primary evidence before contracting a channel, hiring a sales team, or tooling the product. The majority did it in reverse and called the validation step a launch. Industrial markets raise the price of that reversal: a wrong bet in web software burned a funding round, while in industrial AI it burns network capital, OT integration work, and frontline trust that is slow to rebuild.
One wrong hypothesis, six functions paying for it
A wrong commercial hypothesis rarely produces one clear failure. It produces six symptoms, owned by six functions, and from inside none of them look related.
Each team can describe its own problem in detail, and each description sounds like an execution issue local to that team. So the organization responds locally, with sales training, more support headcount, a revised forecast, and an integration task force. None of those mitigations is aimed at the cause, so they keep recurring, because a commercial hypothesis nobody validated regenerates the symptoms as fast as they are managed.
Why the cause goes unnamed
No single function can see all six symptoms, so nobody is placed to recognize one pattern. The people with the broadest view, the initiative's leadership, have the least incentive to name a cause that indicts the original decision.
Tracing six symptoms back to one cause means working across all six functions at the same time. That calls for a structural position more than analytical skill: a view into every function, held by someone independent of the initiative's outcome.
Sales usually shows the pattern first, as conversion below forecast. Our paper on the industrial buying committee sets out how to tell which of its causes is at work.
Why the inside cannot call it
No internal person can credibly challenge the direction of a committed initiative without paying for it in standing, relationships, or career, and that cost shapes what they are willing to say out loud. The team members closest to the evidence also have the most conviction invested in it.
Nobody's integrity is at fault. The incentives everyone in the room works under produce this outcome predictably. The team that spent twelve months building a hypothesis cannot credibly issue a Stop against it, and the sponsor who advocated for the budget will struggle to present the evidence that it should not be spent.
Public commitment raises the price further. Once leadership has told the organization, a board, or a customer that an initiative is working, a Stop costs credibility on top of the capital already spent, and the decision to keep going gets made on reputation.
The same conditions that make an honest internal answer impossible at the point of scaling also raise what the decision is worth. More capital, more headcount, and a deeper market commitment ride on it, so being wrong costs more there than at any earlier point. An independent answer is worth the most at exactly the moment teams most often skip it.
What independence requires
Honest assessment needs three conditions that an internal team cannot create for itself:
- No employment relationship with the client, so the assessment cannot affect anyone's pay or career.
- No equity or success stake in the initiative, so the assessor gains nothing from a yes.
- A written evidence record delivered before the decision meeting, so the political dynamics of the room cannot revise it.
Physical separation is absent from that list. An embedded advisor can sit inside a team every day and remain independent, since independence comes from incentives and has nothing to do with seating. A verdict from someone with no stake in the outcome, and no part in the original announcement, is also easier for leadership to act on publicly than one they would have to originate themselves.
Testing this against your own decision? Request a Sprint Assessment
Frontline evidence decides adoption
In our operator engagements, a maintenance technician's or line operator's unprompted reaction to an initiative is one of the strongest predictors of adoption at network scale. It is most telling before they have heard the management framing or sensed what answer is expected, which is why an operator sprint interviews frontline workers and line supervisors before it interviews management. Management enthusiasm without frontline buy-in predicts a failed rollout just as reliably. The people who will run the system every day know things about their workflow, their equipment, and their incentives that no business case captures.
The workforce data points the same way. In KPMG's 2026 survey of 258 industrial manufacturing technology leaders across 22 countries and territories, 89% agreed that managing AI agents will become a critical workplace skill within five years. Systems are moving closer to the frontline, and adoption gets decided there.
Interview conditions we do not relax
With a manager in the room, what you gather is a performance of alignment, shaped by what is safe to say in front of the person who writes the schedule. People are behaving predictably in an unequal room, and nobody is lying.
So the method requires independent access: conversations at several sites, without management present, scored on the same rubric as every other interview. Problem confirmation needs separate confirmation from the frontline and from management, and management confirmation alone rates Benchmarked. Solution fit is scored separately for each group, so enthusiasm at the top cannot average away skepticism on the floor.
A strong frontline signal gives the deployment case the one piece of evidence most network business cases lack. When the frontline signal is weak, change management readiness rates Not met. Found before the capital commitment, that is a solvable problem; found in month nine of a stalled rollout, it is an expensive one.
Declare the criteria before the research
Purgatory survives on criteria that were never written down. A pilot framed as "let's see what it can do" cannot fail, because success was never defined narrowly enough to be disconfirmed. That makes it a demonstration, and demonstrations rarely produce evidence a capital committee can act on.
A hypothesis has to be able to fail
"Manufacturers want better predictive maintenance" cannot fail. It is vague enough to be true of nearly every manufacturer at some price, in some form, eventually. No interview result could refute it, which means no interview result can meaningfully support it.
"Plant managers at discrete manufacturers with three or more sites will pay a stated annual contract value for this capability, approved through the operations budget" can fail on four separate specifics: the persona, the segment, the price, and the budget path. Those specifics are what make it testable.
Rising interest makes the discipline more important. When nearly every manufacturer is exploring AI, the generic version of almost any industrial AI hypothesis will meet polite agreement everywhere.
What Week 0 produces
Before any research begins, a sprint produces four artifacts, each signed off:
- The hypothesis, written in falsifiable form.
- The pass and fail criteria: the evidence thresholds that will count as confirmation or refutation.
- A definition of the ideal customer profile precise enough to build a target account list against.
- The access path: whether the named accounts come from the client's network or from cold sourcing, which sets the timeline.
A client who arrives with a well-formed hypothesis can compress Week 0 considerably. Most do not, and in our experience the time spent sharpening the question does more for the result than any other stage of the sprint.
Success defined after the fact always matches what was found. If the criteria can move after the interviews, the sprint becomes an expensive way to confirm a decision already made. Declaring them first is what separates a validation sprint from open-ended discovery, and it is what lets the final verdict survive scrutiny from a board or an investor.
Rating every claim
Each claim in the evidence carries a rating:
- Verified: confirmed by scored interviews with real buyers in the specific ideal customer profile, with confirmed urgency, willingness to pay at a specific price, and decision authority.
- Benchmarked: supported by secondary research, such as benchmarks, analyst data, expert interviews, or a technical champion who cannot sign a purchase order. Directionally sound and useful as context, but not enough to base a capital commitment on.
- Assumption: a claim taken as given, stated openly, with a risk statement and a validation path attached.
Organizations overweight Benchmarked evidence because it arrives well formatted. Analyst charts look authoritative and interview notes look messy. Add a confident sponsor and months of momentum, and the difference between "the category is real" and "our buyer will pay" drops out of the conversation. By this definition, every survey figure earlier in this paper is Benchmarked.
The standard allows decisions on claims that are short of Verified, since requiring Verified evidence everywhere would prevent all action. What it requires is that every claim carry its true rating. A Benchmarked claim honestly labeled is a manageable, priceable risk; presented as Verified, the same claim is how a defensible-looking business case becomes a write-off.
Six dimensions against one gut check
A go or no-go meeting that ends in "not quite there yet" triggers nothing. A scored matrix names what failed.
The Decision Matrix scores six dimensions separately, each rated Met, Mixed, or Not met, with the interview citations behind the rating:
| Dimension | The question |
|---|---|
| Problem confirmation | Is the problem real and urgent for the buyer with budget authority, confirmed by that buyer? |
| Willingness to pay | At a specific price, with evidence from solution-test interviews? |
| Market size | Is the confirmed segment large enough to matter? |
| Decision maker access | Can the initiative reach the economic buyer, or only the champion? |
| Competition | What does the buyer compare this to, in their own words? |
| Solution fit | Does the offering match what the buyer described, or what the roadmap assumed? |
Scoring separately prevents the most common error, where a spectacular willingness-to-pay finding hides the fact that nobody interviewed could approve a purchase. A gut check averages those two findings together; the matrix keeps the weak one visible.
Five or six dimensions Met points to Proceed, three or four to Pivot, one or two to Reset, and none to Stop.
Because each dimension is scored against criteria declared before the research, the recommended verdict of Proceed, Pivot, Reset, or Stop follows from the six scores, with the reasoning stated. A decision maker who was not in the room can defend it, and the advocacy that bends most internal go or no-go meetings has little to work with.
It also makes a Pivot actionable. A pricing Pivot, an ideal customer profile Pivot, and a channel Pivot are separate projects, and the matrix says which one to start.
The six dimensions are the core, extended where the situation calls for it. Operator sprints add change management readiness, and sprints for AI vendors add trust, budget architecture, and integration fit. Every question that could kill the commitment gets its own score, its own evidence, and its own line in the findings.
Five fits, assessed together
The matrix sits inside a wider frame. An industrial initiative needs five kinds of fit, and strength in four does not compensate for failure in the fifth:
- Product-market fit: the problem is urgent for the buyer with budget authority.
- Model-market fit: buyers will pay a price that supports viable unit economics, through a budget process they can approve.
- Ecosystem-market fit: the initiative can reach buyers through the channels that control industrial purchasing, such as distributors, system integrators, and OEM relationships, as well as direct sales.
- Team readiness: the team has the domain credibility each buyer touchpoint requires.
- Budget readiness: the runway covers validation before the next capital commitment forces a decision.
A validated product with no channel to the economic buyer cannot currently be sold, and one priced against the wrong budget stalls in procurement. Sold by a team without domain credibility, it loses to a worse product from a trusted one.
The fits are assessed together because they interact. The channel sets the viable price, the price defines the addressable buyer, and the buyer decides how much credibility the team needs. Strength on the other four decides only which function pays for the fifth failure later. The scorecard rates each fit Met, Mixed, or Not met, with a one-line evidence headline and the grade of the evidence behind it, so a decision maker who never sat in the research can read it in two minutes.
The Stop verdict
Of the four verdicts, organizations resist Stop the most, and it often has the highest return.
The arithmetic is plain. An abandoned initiative costs its direct spend, and then opportunity cost, damaged channel relationships, and the organizational credibility spent defending it. A validation sprint that produces a defensible Stop before that commitment costs a small fraction of it. No Proceed verdict, however welcome, protects that much capital.
In our reading, Gartner's projection that more than 40% of agentic AI projects will be canceled by the end of 2027 amounts to a forecast of Stop verdicts issued years late, at post-commitment prices.
What a defensible Stop contains
A defensible Stop is a documented finding, never an opinion that the initiative is bad. It has five parts:
- The specific hypothesis that was tested.
- The pass and fail criteria, declared before the research began.
- The scored primary evidence from the buyers who would have needed to say yes.
- The specific dimensions on which the evidence failed.
- What the research learned about adjacent problems.
That document is the organizational record that makes not committing further capital defensible in front of a board, a leadership team, or an investor.
The fifth part is worth more than it looks. A Stop on one hypothesis is often the sharpest brief for the next, because the buyer who lacked urgency for the tested problem frequently named a problem that does have it.
Stop has to be possible before the sprint starts
If stopping is not a real option before the research begins, because the initiative has already been announced internally as a success, no evidence will produce an honest verdict of any kind. A sprint is a fixed-scope engagement with a defined end, so nobody has a reason to manufacture a path to yes, but that helps only if the organization can act on a Stop. Where Stop was never possible, a Proceed from the same process tells the organization very little.
One standard across three moments
Pilot purgatory is one moment in an initiative's life, and the same standard applies to each of the three.
- Before the commitment, being wrong costs the least and evidence quality counts for the most. It is also when organizations are most confident and least inclined to test.
- When an initiative is not delivering, the question is whether the hypothesis was wrong or the execution was. Those need entirely different responses, and treating them as one problem is how an organization spends eighteen months addressing neither.
- At the inflection before scaling, being wrong costs the most, and internal teams are least equipped to produce an honest answer.
The inflection is the hardest of the three to bring evidence to. By then the program has momentum: parts of it work, capital and headcount are committed, and leadership has said yes in public. That is why it gets the full standard, with nothing lightened.
The discipline is identical at each moment: a falsifiable hypothesis, primary research with real buyers, every claim carrying its true rating, and a named verdict. With one standard across all three, evidence stays comparable over an initiative's whole life. An organization that validated with a scorecard at the first moment can reopen it at the third and see which claims survived and which drifted.
Across a portfolio
The same logic applies across bets. Organizations running many AI and digital initiatives at once often have no consistent evidence standard across them, so resources follow whichever team advocates best. A status update measures activity against a plan and tells you whether the team is executing. Whether the bet should exist is beyond it, because the plan itself embeds the unvalidated hypothesis. Executing well against a wrong hypothesis still ends in a loss.
When every initiative is scored on the same rubric, the bets become comparable for the first time. A portfolio view then shows what no single sprint can: where evidence from different initiatives conflicts, where two teams are chasing the same buyer with different hypotheses, and where the strongest evidence sits in the least-funded area.
The format is one page on purpose. Each area gets one verdict, a board member can read the page in two minutes, and the evidence chain behind each verdict sits in an appendix.
If a scorecard rounds a Stop up to a Pivot to keep the room comfortable, it has stopped being a decision instrument. Where the evidence says stop, the scorecard should say so.
Appendix: diagnosing a pilot that succeeded and did not convert
What the pilot proved
- What exactly did the pilot confirm, and under what conditions?
- Which scale questions were in its scope, and which were assumed?
- What support did the pilot site receive that a network deployment will not?
The decision
- Were pass and fail criteria written down before the pilot began? Where are they?
- Who owns the decision to scale or stop, and when did they last convene it?
- How many times has the pilot been extended, expanded, or handed to a new sponsor?
- Has anyone stated, in writing, what evidence would justify stopping?
The buyer and the budget
- Has the economic buyer for network deployment been interviewed, separately from the champion?
- Does the budget that funded the pilot have authority over the network decision, and has the plant controller checked the value case?
- Does the problem exist at the same urgency at sites without a champion?
The frontline
- Were machine operators and technicians interviewed without management present, at more than one site?
- Do frontline and management reactions differ, and how was each scored?
The distributed cost
- What symptom is each of sales, support, finance, operations, legal, and channel managing?
- Could those symptoms share one upstream cause?
- What did the accounts or sites that stalled say?
The verdict
- Is Stop a real option, or has the initiative already been announced as a success?
- Who would deliver the verdict, and do they have any stake in the answer?
References
- Deloitte, "2025 Smart Manufacturing and Operations Survey," published 1 May 2025: 600 executives at companies with annual revenue of US$500 million or more and US headquarters or operations, surveyed August to September 2024. https://www.deloitte.com/us/en/insights/industry/manufacturing/2025-smart-manufacturing-survey.html Also released through PR Newswire. https://www.prnewswire.com/news-releases/deloitte-survey-reveals-smart-manufacturing-is-driving-advantage-but-needs-focused-investment-and-implementation-302443462.html
- Redwood Software, "Manufacturing AI and Automation Outlook 2026," global survey of 300 manufacturing professionals by Leger, press release, 20 January 2026. https://www.redwood.com/press-releases/manufacturing-ai-and-automation-outlook-2026-98-of-manufacturers-exploring-ai-but-only-20-fully-prepared/
- Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari, MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025," preliminary findings, July 2025 (research period January to June 2025). https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf Also reported by Virtualization Review, 19 August 2025. https://virtualizationreview.com/articles/2025/08/19/mit-report-finds-most-ai-business-investments-fail-reveals-genai-divide.aspx
- S&P Global Market Intelligence, "Voice of the Enterprise: AI & Machine Learning, Use Cases 2025," more than 1,000 respondents in North America and Europe, fielded in late 2024, as reported by CIO Dive, 14 March 2025: "The share of companies abandoning most of their AI initiatives jumped to 42%, up from 17% last year." https://www.ciodive.com/news/AI-project-fail-data-SPGlobal/742590/ Publisher summary: https://www.spglobal.com/market-intelligence/en/news-insights/research/ai-experiences-rapid-adoption-but-with-mixed-outcomes-highlights-from-vote-ai-machine-learning
- Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- CB Insights, "Why Startups Fail: Top Reasons," 5 March 2026, covering 431 venture-backed companies that shut down since 2023, 385 of them with identifiable failure reasons. https://www.cbinsights.com/research/report/startup-failure-reasons-top/
- CB Insights, "The Top 20 Reasons Startups Fail," 2014, analysis of 101 startup failure post-mortems, in which startups could cite more than one reason. https://s3-us-west-2.amazonaws.com/cbi-content/research-reports/The-20-Reasons-Startups-Fail.pdf Also reported by Inc. https://www.inc.com/erik-sherman/42-percent-of-startups-that-fail-trip-over-this-one-critical-avoidable-point.html
- KPMG, "Global Tech Report 2026: Industrial manufacturing," 258 industrial manufacturing technology leaders across 22 countries and territories. https://kpmg.com/xx/en/our-insights/ai-and-technology/global-tech-report/industrial-manufacturing.html
What is not yet sourced
- The S&P Global figures (reference 4) are cited through CIO Dive's reporting, which quotes the 42% and the 17% a year earlier. S&P Global's own pages could not be retrieved when checked, and the underlying report has not been reviewed.
- The MIT NANDA figure (reference 3) has been checked against the report. The authors describe the report as preliminary and its figures as directionally accurate, based on interviews and a survey with the limits they list.
- The CB Insights figures (references 6 and 7) have been checked against both reports. The 2014 count includes every reason cited and the 2026 count a primary reason, which is why the two are not a time series.
- The opening observation that many pilots succeed and then stall rests on Thing Company's engagement experience. The figures in the section on the numbers show the same gap from outside, but none of them measures it directly.
- The predictive value of frontline reactions, the six-function cost pattern, and the portfolio observations rest on Thing Company's engagement experience, not a published study.
- The Decision Matrix, the five fits, the Week 0 artifacts, and the evidence ratings are Thing Company methodology.
- The recurrence of the sequencing pattern across technology transitions, its spread to agentic and physical AI, and the relative value of Week 0 are Thing Company's reading from engagement experience, not a measured result.
About Thing Company
Thing Company is an independent market validation practice for industrial technology. We test whether a buyer exists at a price that works.
When a pilot has succeeded and not converted, the work is to convene the decision it has been waiting for. How we work sets out the evidence standard, and the Sprint is the instrument that applies it.
Holding a pilot that will not convert? Start with a Sprint Assessment.
Not ready to talk? Run the free stall diagnosis self-check from the toolkit.