THING COMPANY PAPERS
Governing agents in production: what to ask before an agent acts
An independent evaluator's paper on governing AI agents, built from public sources: where approval gates belong, why autonomy is set per task from outside the agent, why controls sit on the dispatch path, and the 31 questions to put to a vendor.
Harinderpal Hanspal · LinkedIn · hans@thing.company · About 23 min read · 10 sections · Appendix · References
Request a Sprint Assessment A thirty-minute call about your own decision, not a purchase.
Executive summary
Most writing on agentic AI governance comes from people describing what they sell. This paper comes from an evaluator's seat. Thing Company scores other companies' agentic systems for buyers and investors, and the argument below is the reasoning we bring to that work. It rests on public sources (regulators' texts, government guidance, surveys, vendor documentation, reported incidents), each cited with its scope and limits. Where no source covers a point, we picture a plant scenario and say so.
The argument is short.
Govern an agent by the damage a wrong action can do, and by where that damage leaves your control. We call this blast radius.
Four rules follow. The first puts people on the outward boundary. An agent that is wrong internally costs a regeneration, while one that is wrong toward a supplier, a customer or a machine costs something you cannot get back, so the approval gate belongs on that edge and has to be cheap enough that nobody routes around it. The second sets autonomy per task, from outside the agent: a graduated scale lets a documentation fix and a setpoint change carry different levels where a single supervised-or-autonomous switch forces one trust decision onto everything. The third enforces controls below the agent, on the path every action takes, where a prompt cannot argue them away. The fourth matches verification depth to the cost of silent failure, because components that each pass their own tests can still compose into a path that fails.
Two pairs of capabilities get confused. A tamper-evident audit trail proves a record was never altered, which is a different job from proving who wrote it. A spend cap keeps the bill under a ceiling, which is a different job from showing what the money earned.
The evidence is thinner than vendor pitches suggest: we found no standards body that has published autonomy levels for AI agents, and no regulator that requires tamper-evident logs.
The appendix lists the 31 questions we use when we score an agentic system, grouped by the five dimensions they evidence.
Why this paper reads from the outside
We do not describe our own systems here. Two facts about our practice are enough: we use agents in our own research and outreach, and a person approves anything that leaves the building.
Every figure carries its publisher, date, sample and scope, because a number stripped of its scope is how bad claims travel. A survey of 600 breached organizations is not a survey of all organizations. An approval rate is not an error rate. A prediction is not a count. Scenarios that begin "picture" are hypothetical, built for an industrial reader, and carry no data.
Blast radius sets the gate
Picture an agent working for a plant's purchasing team. It reads open purchase orders, spots a late delivery, drafts a note to the supplier and proposes an expedite request with a freight charge. The first three steps stay inside the building, and a wrong guess costs a regeneration. The expedite request, once sent, commits the plant to a charge and tells a supplier something that cannot be unsent. Generation was never the hard decision. Release is.
Two ways to put the gate in the wrong place
Gate everything, and approval becomes a formality. Anthropic reported in March 2026 that Claude Code users approve 93% of permission prompts, and built classifiers to reduce approval fatigue (Anthropic Engineering, 25 March 2026). That is an approval rate for one product, not a measure of safety. A gate that always gets a yes adds delay and no review.
Gate nothing, and one wrong write costs more than any queue would have. In July 2025 a coding agent deleted a production database during a declared code and action freeze, and the company's chief executive called it unacceptable and said it should never be possible (Fortune, 23 July 2025). It is one reported incident, not a rate.
OWASP names the middle path. Excessive agency is risk LLM06 on its 2025 list, arising "whether due to hallucination or prompt injection attacks", and the advice is human-in-the-loop control "to require a human to approve high-impact actions before they are taken" (OWASP, LLM06:2025). High-impact is the word that carries the weight. The agent works alone inside the boundary, and a person decides what crosses.
What makes a gate work
Three questions settle the design.
Where does the reviewer already work? A pending line on the screen a buyer already uses gets a real decision, while a separate approvals queue gets visited once a day and cleared in a batch.
Does the approver see the request or the result? "Send the expedite note" tells a person nothing. The card should show the supplier, the charge and what the order said before.
What does approval cost? If reviewing takes longer than regenerating, people route around the gate, and an unofficial bypass is worse than no gate because the record of what was approved becomes false. For high-risk systems, Article 14 of the EU AI Act asks that overseers stay aware of "automation bias" and be able to override, reverse or stop the system, as published on artificialintelligenceact.eu. It applies to high-risk systems, not to every agent.
Also ask whether an item approved from a second place, such as an inbox, resumes the paused run. If not, the gate looks closed and is not.
What the line costs
Throughput pays for it, since a person approving every outward message moves at the speed of their attention. Deloitte's 2026 survey of 3,235 senior leaders in 24 countries, fielded August to September 2025, found that only one in five companies has a mature model for governing autonomous agents (Deloitte, State of AI in the Enterprise 2026). The arithmetic still favors the line: a supervised send costs minutes, and an unsupervised wrong one costs a supplier relationship. If a vendor says a person on the boundary costs nothing, ask how they counted.
Autonomy is a dial, set from outside the agent
Picture a plant's scheduling agent on one afternoon. It fixes a typo in a shift-handover template, and it adjusts the batch sequence sent to a line. Suppose the tool offers one setting for both. Turn autonomy on and the sequence change ships unreviewed. Turn it off and the typo waits a day for a person to confirm something that carried no risk. A tool with one switch has not made a careful choice. It cannot tell a typo from a production change, so it hands that judgment to a person every time.
Levels per task
A graduated scale asks a sharper question for each task. May the agent read but not act? Propose, with a person confirming? Act within stated limits and report afterward? Act with no gate? The answer can differ between two jobs the same agent does five minutes apart.
The field offers proposals, not a standard. Three researchers proposed five levels in June 2025, defined by the role the user plays: operator, collaborator, consultant, approver and observer (Feng, McDonald and Zhang, arXiv 2506.12469). NIST announced an AI Agent Standards Initiative on 17 February 2026, with pillars on standards, open protocols and research on agent security and identity (NIST); it does not define levels. Singapore's IMDA launched a voluntary framework for agentic AI on 22 January 2026 (IMDA); law-firm summaries say it asks organizations to set limits on an agent's autonomy, and we have not read its text.
So a vendor's "autonomy levels" are its own scale. Ask where the list is defined, and whether the same list appears on every screen and in every log.
Only an operator raises the level
If an agent can raise its own autonomy, the scale is decoration. A person outside the agent sets the level, and a workflow that earns more trust gets it by someone changing the setting. The OWASP text explains why: excessive agency can arise through prompt injection, so a limit written into the agent's instructions is one the agent can be talked out of. A limit applied below the agent cannot.
Something should also lower the level without a person. CISA and eight partner agencies advise that where AI actively updates control logic, systems should add human-in-the-loop intervention points, because problems "can escalate before operators become aware of them" (CISA and partners, 3 December 2025). It is guidance, not regulation. Ask what lowers a level after repeated failures, and whether that rule has ever fired in production. A rule with no caller protects nobody.
Controls belong below the agent
A control works when it sits on the path every action must take and is enforced by the layer the agent runs on. Beside that path, or inside the agent's own code, it fails in ways you can predict.
Policy before the action
Every consequential action should pass a policy check before it happens, run by the platform and not by the agent. A guardrail the agent implements for itself is one a bad enough prompt can argue past. Every request should carry a label for the class of data it touches, checked at routing time.
Routing is a data decision
The leak is measured. In IBM's 2025 study of 600 breached organizations (March 2024 to February 2025), one in five reported a breach involving shadow AI, meaning AI tools staff used without approval, and only 37% had policies to manage or detect it. Organizations with high shadow-AI use had breach costs $670,000 higher than those with low or none, a difference between two groups and not the price of one breach (IBM, 30 July 2025).
No primary source we found says "send this class of data to a local model and that class to a hosted one." What exists is principle. NIST's generative AI profile warns that models "may leak, generate, or correctly infer sensitive information about individuals" and suggests documenting reliance on third parties and identifying fallbacks (NIST AI 600-1). Turning that into a routing rule is policy work, and it belongs in the layer every call passes through. It means deciding three things:
- Which providers exist at all. This should be an allowlist, not API keys scattered through configuration files.
- Which data classes may reach which providers, checked at dispatch for each request.
- What happens when the preferred model is unavailable. A fallback chosen during an outage inherits none of the review the primary received.
Models also disappear. Anthropic commits to at least 60 days' notice before retiring a publicly released model, and notified one retirement on 5 June 2026 for 5 August 2026 (Anthropic, model deprecations). A swap is not a drop-in, so whoever approves one should re-check the data rule, not only output quality.
A fast test for any system: search the code for provider SDKs. If a component imports one outside the routing layer, the architecture diagram is wrong about where data can go.
Consent on the dispatch path
A contact unsubscribes on Tuesday afternoon. On Wednesday morning an outbound agent working from Monday's list sends a follow-up. Nobody wrote code to ignore the request; consent lived in a system the agent never checked.
The duty attaches to the sender, whatever pressed send. The FTC's CAN-SPAM guide says to honor an opt-out within 10 business days, makes no exception for business-to-business email, says a sender cannot contract away legal responsibility, and puts a penalty of up to $53,088 on each violating email, as adjusted in the guide (FTC, August 2023, edited January 2024). Canada's anti-spam law reaches anyone who causes or permits a commercial message to be sent without consent (Justice Laws Website, S.C. 2010, c. 23), a phrase an agent operator should read twice. GDPR Article 21 lets a person object to direct marketing at any time and names no grace period (gdpr-info.eu).
A next-day message does not breach the US window on its own. The operating problem exists anyway: an agent sends far more messages than a person in the same time, so the gap between a consent change and a message that ignores it narrows as send rates rise. The check has to run at each send, on the same path as the autonomy gate and the budget check, because a review after the campaign can only report what already went wrong.
Passing tests do not show that a check refuses anything: it can run on every send and block nothing if it reads a field no sender fills in. So probe it. Put an address you control on the suppression list and send to it through every route that reaches a mailbox, including a manual send, a retry and a second channel. Expect a refusal from each. Then make the suppression source unreachable and send again. A safe system skips the send, and one that fails open mails everyone.
Testing this against your own decision? Request a Sprint Assessment
Liability follows actuation
On one side of an integration, a wrong agent action is a bad database row. On the other, it is a setpoint on a machine. The questions that reach lawyers and regulators ask what touched the machine, under whose authority, and with what checks.
What the guidance says
CISA and eight partner agencies published "Principles for the Secure Integration of Artificial Intelligence in Operational Technology" on 3 December 2025. It is guidance for critical-infrastructure owners and operators, and it covers AI agents. Four statements matter to a buyer (CISA and partners):
- Language-model AI "almost certainly should not be used to make safety decisions" in OT environments.
- Where AI actively updates control logic, add human-in-the-loop intervention points.
- Build failsafe mechanisms, and include how to bypass or replace the AI system in functional safety and incident response processes.
- Prefer push-based or brokered architectures that give the AI no persistent inbound access to the OT network.
Why the check cannot live in the device
NIST's OT security guide says many field-level devices and protocols, including sensors and actuators, cannot be authenticated, so data can be replayed, modified or spoofed (NIST SP 800-82 Rev. 3, section 5.3.6). The guide does not discuss agents. The consequence for them is ours: a device that accepts a command from anyone on the network cannot tell an operator from an agent from an attacker.
Records show what unauthenticated write access costs. A 2025 CVE for three Advantech I/O modules describes an unauthenticated attacker sending Modbus TCP packets to manipulate digital outputs (CVE-2025-48466); it concerns those modules, not every Modbus device. In advisory AA23-335A, actors reached internet-connected Unitronics controllers with a default password or none, replaced control logic and disabled critical functions, including in water and wastewater facilities (CISA and partners). Neither involved an AI agent. They show the layer an agent would write to.
So the control moves upstream. Whatever decides that an agent may write has to run before the write, in software that knows who is asking, what the action does and who approved it. The protocol gateway is the last point where every action is still a request that can be classified, gated, logged and refused. Three questions test it. Is the list of what an agent may do built when it asks, so a revoked capability vanishes from the next answer, or is it a document nothing enforces? Is each action classified by consequence before it runs, and not only by its tool's name? Does every route to the device pass the check? A vendor who answers the last with "the policy engine" should be asked to show every route.
The commercial side
Once a system can move equipment, the buyer's health, safety and environment function and its legal counsel join the purchase, and either can stop it. Under the AI Act's revised dates, high-risk obligations apply from 2 December 2027 for stand-alone systems and 2 August 2028 for those embedded in products, and machinery-regulation products are exempted from direct applicability (Council of the EU, 30 June 2026, via EEAS). Whether a given plant system is in scope is a legal question.
Under our evidence standard, "we are working through the liability model" rates as an Assumption. A structure the buyer's own legal and safety teams have reviewed and accepted rates Verified. Ask the buyer who would sign off on the liability terms, then interview that person. Our paper on validating physical AI covers that dimension.
An audit trail that cannot name the actor
Picture a valve setting that changes at 2 a.m. in a plant where four agents share access to one system. The log shows which record changed and by how much, with a signature proving nobody edited the entry afterward. It cannot show which of the four agents acted, on whose authority, or whether the sub-agent that acted was the one its parent believed it had started.
Two capabilities are in play. A hash-chained trail answers whether anyone altered a record after the fact: forge an entry and the chain breaks. Identity answers whether the writer was who it claimed to be, what started it, and whether the authority it used had been delegated.
What the rules ask of a log
Article 12 of the EU AI Act says high-risk systems must technically allow automatic recording of events over the system's lifetime, with logging that supports traceability (Article 12). Article 19 has providers keep the logs for at least six months unless other law says otherwise (Article 19). Both are as published on artificialintelligenceact.eu, because EUR-Lex could not be read.
Read what is absent. Neither article says tamper-evident, immutable or cryptographically chained, and we found no regulator that requires it. Both apply to high-risk systems only, and the Digital Omnibus moved the Annex III obligations to 2 December 2027, according to Hunton. Anyone saying the duty starts in August 2026 is quoting the old schedule. The CISA guidance above separately says to log and monitor the inputs and outputs of AI components.
Three claims a log can make
"Full audit trail" can mean three things. The record is unaltered: hash chaining does this, a design choice and a good one for an investigation. The record was written by the system it names: a signature does this, while the key stays protected. Or the record says which agent acted, who authorized it, and why it was allowed. The first two do not touch the third. A perfectly chained log can show that entry 4,112 is intact and that it came from "the agent service account", shared by all four agents.
OWASP's list of non-human identities (service accounts, roles, keys) does not address agents that start other agents (OWASP Non-Human Identities Top 10), and we found no standard answer to the lineage question. The gap shows in breach data: of 600 breached organizations in IBM's 2025 report, 13% reported a breach of an AI model or application, and 97% of that group lacked proper AI access controls (IBM).
Ask which half the vendor means, in writing. Then pick one action an agent took last week and have the vendor explain it from the log alone. Which agent did it, with an identity unique to that agent? What authorized it? Was the decision recorded when it was made or reconstructed later? If another agent started this one, can you see the parent run? A regulator's first question is who did this, and a trail that survives forensics without naming the actor cannot answer it.
Spend caps and cost ledgers
A cap sits in front of the call, adds each call's spend against a ceiling, and refuses the next call when the ceiling is reached. That is the right shape for stopping a runaway loop. It also produces exactly one number per account.
Picture three agents in a plant's AI workspace: one matching invoices, one chasing late suppliers, one proposing schedules. All three stay inside the monthly cap, and nothing alarms. At the funding review someone asks which of them was worth what it cost, and the cap has no answer because it never recorded who spent what.
What a funding review asks
Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls. The causes are alternatives, and the figure is a prediction built on a January 2025 poll of 3,412 webinar attendees, not a count of cancellations (Gartner, 25 June 2025). It points to one mechanism worth naming: a team with disciplined spend controls can still lose the argument for its project, because staying under budget and demonstrating value are separate capabilities. No alarm fires when attribution is missing, so the failure first shows up at the funding review.
Where caps run out
A cap's estimate is a floor, since a check before the call can price the prompt but not the answer. Its coverage follows the code path, so a second path or a tool that makes its own model calls goes unmetered. And a real cost record gives each event who asked and for what: workspace, workflow, run and the identity behind it. Without that, a report says what was spent and never what it bought.
The local bill needs an owner
Running agents on hardware you own is often argued from privacy and sovereignty, and both are real. The time goes into the bill after deployment, and updates come first. NIST's patch guidance says patching "should be considered a standard cost of doing business", and that delaying routine patching "also makes emergency patching more difficult, time-consuming, and disruptive" (NIST SP 800-40 Rev. 4, April 2022). The clearest picture of a bad update reaching a fleet is not industrial: Microsoft estimated that CrowdStrike's July 2024 update affected 8.5 million Windows devices (Microsoft, 20 July 2024). We found no primary figure for how often updates cause outages on industrial fleets.
Cloud inference produces an invoice, and invoices have a route through an organization. Local inference produces depreciation, power, an update rota and the need for someone who notices a failing machine. A cost tool that prices an owned machine at zero is right about vendor spend and wrong about the cost of serving. A cost nobody owns eventually arrives as a surprise.
Verification depth follows blast radius
Picture a supplier portal at a plant, with a contact form updated using a coding agent. Every page returns 200, the build is green and the form renders. A supplier fills it in and presses submit, and the request is refused because the server's origin check rejects the site's own address. Nothing a status-code sweep reads can see that.
Each check is correct about its own layer. The thing a supplier came to do is a path through every layer, and when separate pieces are generated separately, the seams belong to nobody.
Veracode tested more than 100 models on 80 coding tasks and reported on 30 July 2025 that the generated code introduced an OWASP Top 10 flaw in 45% of cases (Veracode). That is 45% of task attempts in one benchmark, by a security vendor, not 45% of production code. The deeper problem is misjudgment. METR's July 2025 trial of 16 experienced open-source developers on 246 tasks found them 19% slower with AI tools after they predicted a 24% speedup, and afterward they still believed they were 20% faster (METR). METR's February 2026 follow-up changed its design and says true gains are probably higher (METR). That work measured speed, not defects. What carries over is the gap between what people feel and what a measurement shows.
So set verification depth by the cost of a path failing silently, whatever the effort of checking it. Uniform verification tends to be uniformly shallow. An internal report page can live with a status sweep. A public form, a payment, a work order, or anything that writes to equipment needs a run that does what a user does and confirms the outcome: the message arrived, the record exists, the setpoint changed to the intended value. And since a pre-release test does nothing about the failure that gets through, CISA's advice is the other half: a documented way to bypass or replace the AI system.
For a buyer, this is a rule about success criteria. An acceptance test that checks layers and infers that the path works is not falsifiable. Write the criterion as the outcome.
Where governance meets the buyer
Everything above is architecture. Buyers turn it into a commercial question, and they do so earlier than most vendors expect.
With a recommendation, a person absorbs a bad call. When the system acts on a production schedule, a maintenance plan or a supplier order, the buyer trusts it with outcomes they answer for personally, and that trust does not grow with model accuracy. It is granted at a specific level of autonomy, for a specific use case. Gartner's Manufacturing Predicts 2026 report, dated 10 December 2025 and known here through published summaries, expects semiautonomous AI agents to orchestrate 10% of key production, quality and maintenance use cases by 2030, up from 2%, with humans keeping final approval. A buyer on the advisory step will not approve a system priced and packaged for full autonomy, however accurate it is.
We score trust as its own dimension in discovery interviews for that reason. The same plant manager may trust a system to flag an anomaly, hesitate to let it reschedule maintenance, and refuse to let it stop a line, so trust is scored per decision on a one-to-five scale, with a strong signal at an average of 4.0 or above. A low score is a product-configuration finding. It points toward a version where the system recommends and a person confirms, with more autonomy earned later against an operating record.
These mechanisms let a vendor meet the buyer on the step the buyer is on. A graduated scale lets one product run as advisory at one account and autonomous at another. A gate on the outward boundary gives the operations lead a decision they can own. An identity trail and a cost ledger answer what legal and finance will ask before procurement signs. A vendor without them can still demo well, and the deal stalls later in a review it did not know existed.
Seen from the buyer's side, Gartner's three causes for canceled agentic projects are each a validation failure: costs escalate when integration and autonomy assumptions were never tested, value stays unclear when nobody confirmed the outcome with the buyer, and risk controls fall short when nobody scoped the trust threshold or the liability model. The mechanisms here answer the third. The first two need evidence from the buyer.
Appendix: governance readiness questions
These are the questions we use when we score an agentic system, grouped by the dimension they evidence.
Governance readiness
- Do policy checks run before each consequential action, and are they enforced by the layer below the agent or by code inside it?
- Does the system know which class of data each request touches, at routing time?
- Where does the suppression check run: on the dispatch path for each send, or in a review after the campaign?
- How stale can consent data be at the moment of sending, in the worst case? Measure it. Do not accept a statement of intent.
- When an unsubscribe arrives mid-sequence, what stops the next message: code, or a person who has to notice?
- Can anyone alter an audit record after the fact without detection?
- Can the record prove which agent acted, what started it, and under whose delegated authority?
- When the vendor claims a full audit trail, which of questions 6 and 7 are they answering? Get the answer in writing.
Autonomy tolerance
- Where is the outward boundary, and what does the organization pay to keep a person on it?
- Do approval gates sit on that boundary, or are they scattered by feature?
- What does approval cost the approver, and would you make the same decision at that price?
- Is autonomy set per task class, or does one toggle govern everything the agent touches?
- Is the autonomy scale defined once, below the applications, or does each tool define its own?
- Can an operator see the current autonomy level everywhere it applies, and change it without a deploy?
- Who can change an autonomy level: an operator outside the agent, or the agent's own configuration?
Budget architecture and ROI ownership
- Can any component reach a model provider without passing the shared routing layer? Check the imports, not the diagram.
- Does routing policy consider data classes, or only price?
- Who reviewed the fallback chain, and when?
- Is there a spend cap, and is it enforced before dispatch or reported afterward?
- Can the system attribute cost to a specific agent, workspace, and workflow?
- Who can say which agent earned its spend last month, and how long does producing that answer take?
- When an update fails on a machine the vendor cannot reach, whose runbook applies, and who travels?
- Is "fleet" one governed unit or many administered machines, and does the vendor know which claim they are making?
- Where does the interface run ahead of what executes, and is that boundary written down?
- Who in the buying organization owns the local operating bill, by name?
Liability and safety
- Is the capability catalog assembled at dispatch, or is it documentation? Revoke a capability and time how long it takes to disappear.
- Are actions classified by consequence before execution, with reads and writes on different terms, and agents and people on the same terms?
- Does anything lower autonomy automatically after repeated failures, or does the design assume a person is watching?
Falsifiable success criterion
- Does the acceptance test drive the real interface, or does it check layers and infer that the path works?
- Is verification depth assigned by blast radius, or applied uniformly?
- When was the highest-consequence path last exercised end to end by something other than a customer?
References
Each source is linked where it is cited. The ones the argument leans on most:
- Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- IBM, "IBM Report: 13% of Organizations Reported Breaches of AI Models or Applications, 97% of Which Reported Lacking Proper AI Access Controls," 30 July 2025, on the Cost of a Data Breach Report 2025 (Ponemon Institute). https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls
- US Federal Trade Commission, "CAN-SPAM Act: A Compliance Guide for Business." https://www.ftc.gov/business-guidance/resources/can-spam-act-compliance-guide-business
- Gartner, "Manufacturing Predicts 2026: AI Agents, Digital Twins and the Race to Autonomous Operations," 10 December 2025 (paywalled; the figure is confirmed through two independent summaries, Synera's and Bassetti's). https://www.synera.ai/analyst-study/gartner-manufacturing-trends-2026, https://www.bassetti-group.com/en/manufacturing-gartner-2026-ai/
- CISA, NSA, FBI and partner agencies, "Principles for the Secure Integration of Artificial Intelligence in Operational Technology," 3 December 2025. https://www.cisa.gov/sites/default/files/2025-12/joint-guidance-principles-for-the-secure-integration-of-artificial-intelligence-in-operational-technology-508c.pdf
What is not yet sourced
- The Gartner figures (references 1 and 4) were confirmed through secondary summaries, because Gartner's pages blocked automated reads.
- The IBM figures (reference 2) describe 600 organizations that had experienced a breach, so "13%" and "one in five" are shares of that group. They come from the press release; the full report was not reviewed.
- EU AI Act articles are quoted as published on artificialintelligenceact.eu, and the Omnibus dates rest on a law-firm note and a Council release mirrored on the EEAS site. Check the Official Journal before relying on an article number or date.
- Singapore's framework is described through law-firm summaries; we have not read its text.
- Lineage as an unsolved problem is our reading of OWASP's list.
- Consent rules outside the texts cited have not been reviewed, and the EU position varies by country for business recipients.
- Update outages on industrial fleets have no primary rate behind them, and the CrowdStrike event involved Windows hosts.
- The scenarios that begin "picture" are hypothetical and carry no data.
- The 4.0 trust threshold is Thing Company methodology, not an external benchmark.
About Thing Company
Thing Company is an independent market validation practice for industrial technology. We test whether a buyer exists at a price that works.
We use agents in our own research and outreach, and a person approves anything that leaves the building. We score other companies' agentic systems for buyers and investors, which is where the questions in this paper's appendix are used. How we work sets out the evidence standard, and the Sprint is the instrument that applies it.
Building, buying, or funding an agentic system? Start with a Sprint Assessment.