Cost of Running AI on Your Own Hardware
Standards and a 2024 outage show why an update to machines you own is an outage risk, and which cost lines a local AI budget usually leaves blank, with a worksheet to fill in.
By Harinderpal Hanspal on July 2026. Updated October 2026
Most of the cost of running AI on your own hardware arrives after install: updates on machines you have to reach, a runtime to administer on each one, and GPU time that cost accounting often leaves at zero. NIST treats patching as a standard cost of doing business.
Privacy gets argued first when a plant considers running AI on its own machines. The bill that follows is quieter, and it starts with a question nobody budgets for: how do you update a box you have to drive to?
An update is a scheduled outage
NIST's guide to enterprise patch management says business owners may believe patching hurts productivity, because it needs scheduled downtime and adds the risk of more downtime if something goes wrong. Its answer is that "patching should be considered a standard cost of doing business." It also warns that routine patching gets postponed, and that postponing makes the eventual emergency patch harder and more disruptive. The guide covers IT, OT and IoT (NIST SP 800-40 Rev. 4, April 2022). It is generic guidance, not an OT outage statistic. NIST names the cost anyway.
A local model runtime adds to that load. The model server, its drivers and the models themselves each have versions, and each machine runs its own copy.
What a bad update does at fleet scale
On 19 July 2024 a CrowdStrike content update affected about 8.5 million Windows devices, which Microsoft put at under one percent of all Windows machines (Microsoft, 20 July 2024). CrowdStrike's root cause analysis says the sensor expected 20 input fields and the update supplied 21, which caused an out-of-bounds memory read and a crash. Its remedy was a staggered deployment starting with a canary, and more customer control over when and where updates land (CrowdStrike, August 2024).
To be clear, that was a security vendor's content update on Windows hosts, not an industrial controller and not an AI runtime. What carries over is the mechanism: one bad push, many machines at once, and recovery that needs someone physically or remotely at each. The fix the vendor chose was staging.
For electric utilities the rule is written down. NERC CIP-007-6 requires evaluating security patches for applicability at least every 35 calendar days, then applying the patch or creating a dated mitigation plan within 35 days of the evaluation, for medium and high impact systems (NERC CIP-007-6, parts 2.2 and 2.3). It is not "patch within 35 days", and which version is enforceable is worth checking before you cite it. Most plants face no such rule, and that is the problem: nobody is scheduled to do the work.
Where a local AI budget goes blank
A cloud invoice hides the update burden, because the provider carries it. On hardware you own, a failed update is your outage on your own floor, and sometimes a drive to wherever the box sits. Three habits leave the line empty.
The first is counting only purchase price. NVIDIA's launch price for the Jetson AGX Thor developer kit was $3,499, with 128 GB of memory (NVIDIA, 25 August 2025). That is a launch price for a developer kit. It says nothing about what the box costs to keep current for five years.
The second is administering each machine as its own unit. Ten machines with ten runtimes means ten update decisions, ten health checks and ten places for a model file to differ. Staged rollouts, with a canary first and a rollback, are the answer the CrowdStrike analysis points to, and a small site rarely has them for AI runtimes.
The third is the GPU line. Cost accounting built around invoices has nothing to record when the machine is owned, so a local model call tends to book at zero, and a GPU sitting warm with a model loaded costs the same as one working. The expensive thing on a fleet is often idle accelerator memory, which token accounting cannot see.
Memory sizing belongs here too, since buying a machine that cannot hold the model at your context length is a cost with no recovery. The memory arithmetic shows the sum, and the break-even note shows how these lines feed a per-token figure.
A worksheet of local AI cost lines to fill in
| Line | Question to answer | Your number |
|---|---|---|
| Hardware | Purchase price over useful life, with the replacement date | |
| Power | Measured wall power under load times your rate | |
| Idle GPU time | Hours a model sits loaded with no work | |
| Update labor | Hours per machine per update, times updates per year | |
| Reach | Can you update remotely, or does someone drive there? | |
| Rollback | What happens to the line if an update fails? | |
| Ownership | Whose budget carries the bill, by name? |
A cost nobody owns turns up later as a surprise. Put a name on the last row first.
NIST SP 800-40 Rev. 4 (April 2022), NERC CIP-007-6, Microsoft's post of 20 July 2024, CrowdStrike's root cause analysis (August 2024) and NVIDIA's launch release (25 August 2025), all read on 6 October 2026. No figure here is a measurement of any deployment.
Related notes
- Open or closed weights: settle the data policy before the model
- AI agent cost tracking: a spend cap cannot say which agent earned its bill
Related insights
Related paper: Combining agentic AI with distributed orchestration at the industrial edge