Local Model vs API Cost: What to Measure
A break-even between a local model and a paid API is one division. Only the API side has published numbers, as of 6 October 2026, so here is the sum, the prices and what to measure.
By Harinderpal Hanspal on October 2026. Updated October 2026
Local model vs API cost is one division: the monthly cost of the local machine over the API price per token gives the volume where they match. Only the API side has published numbers. No primary source gives a cost per token for local hardware.
Every "local is cheaper" claim hides a division. Monthly cost of running the model on your own machine, divided by the API's price per token, gives the monthly volume at which the two cost the same. Below that volume the API wins. Above it, local wins, if the machine can actually serve that volume.
That is the whole method. The trouble is that only one side of it has published numbers.
What the API side looks like on 6 October 2026
Prices per million tokens, input and output, from each vendor's own page as of 6 October 2026:
| API | Input | Output |
|---|---|---|
| Claude Sonnet 5.5 (Anthropic) | $2 | $10 |
| Claude Haiku 4.5 (same page) | $1 | $5 |
| OpenAI gpt-6.1-sol (OpenAI) | $2.00 | $10.00 |
| OpenAI gpt-6-luna (same page) | $0.10 | $0.50 |
| Together AI gpt-oss-120B (Together AI, undated page, read 5 October 2026) | $0.15 | $0.60 |
Anthropic's batch API takes 50% off both prices, and its US-only inference option carries a 1.1x multiplier. OpenAI's batch is also 50% off. Prices move: Google lists Gemini 3.8 Flash at $0.75 input through 31 December 2026, then $1.50 (Google). A break-even worked out today expires with the price.
The last row of the table deserves attention. When the open-weight model you would run locally is also sold per token by a host, that price is your real competitor, and it sits well below the frontier closed models.
Why a price per token does not compare
Anthropic states that Claude 4.7 and later models use a tokenizer that produces approximately 30% more tokens for the same text, and tells readers to recount prompts against the model they plan to use (Anthropic pricing page). The same sentence at the same listed price therefore costs more on the newer model. We found no documented tokenizer ratio between OpenAI, Google and Anthropic, so we state none.
Compare cost per task, with token counts measured on each model. A cost per token tells you little until you know how many tokens your task takes.
The five inputs on the local side
A local machine costs the same loaded and idle, and agent work arrives in bursts. So the denominator is the volume the machine serves, not the volume it could.
- Hardware. Purchase price over useful life. NVIDIA's launch price for the Jetson AGX Thor developer kit was $3,499 (NVIDIA, 25 August 2025), a launch price for one product.
- Power. Measured wall power under load, times your rate.
- Operations time. Updates, a runtime per machine, someone who notices when it stops. The cost note on hardware you own lists the lines.
- Utilization. Served tokens divided by what the machine could serve.
- Measured throughput. Tokens per second for your exact model and quantization, on your exact box.
We found no primary source for the last three on any given machine. Cost accounting built around invoices tends to treat an owned machine as free, since it never sends one. A comparison that books local at zero will always find local cheaper, and that proves only the bookkeeping.
The measurements to take before anyone quotes a number
| Measure | How | Month of real use |
|---|---|---|
| Output and input tokens served | Count at the server, by task type | |
| Wall power | Meter the box under load and idle | |
| Operations hours | Log time on updates and incidents | |
| Hardware cost per month | Price over useful life | |
| Input to output ratio | From your own traffic | |
| Tokens per task on the API model | Recount on the model you would buy |
Input is cheaper than output on every price above, and an agent that reads long context spends mostly on input, so use your own ratio. Divide the sum of the four cost rows by the API price at that ratio. A vendor who hands you a break-even without these inputs is quoting a guess.
API prices from Anthropic's pricing page, OpenAI's pricing page and Anthropic's tokenizer note, read on 6 October 2026 (Together AI's page is undated, read 5 October 2026). No figure here is a measurement, and we publish no break-even.
Related notes
- The cost of running AI on your own hardware starts with updates you must reach
- Which AI requests may leave the building: sort by data class first
Related paper: Combining agentic AI with distributed orchestration at the industrial edge