LLM Tokens per Second: What a Speed Claim Means

A tokens-per-second figure can mean prompt processing, generation or total time divided by length. How to read the claim, why a non-streamed reply hides the rate, and what to ask for.

By Harinderpal Hanspal on August 2026. Updated October 2026

A tokens-per-second claim is only as good as its definition. Benchmark tools report prompt processing and generation separately, and a rate cannot be read from a reply that arrives in one piece. Ask which of the numbers a vendor measured, and with the reply streaming.

A reply that arrives in one piece gives a wait time, not a rateSketch of two timelines. The top one shows a single large block arriving at the end of a wait, labelled with the wait as measurable and a speed of none in red, because a rate cannot be read from one piece. The bottom one shows many small ticks spread over time, from which a rate can be read. A note says to ask which rate a vendor measured.reply arrives in one piecewait: measurablespeed from it:nonereply streams in piecesticks over time: a rate can be readask which rate the vendor measured
A reply that arrives in one piece gives a wait time, not a rate

IDC's February 2026 edge forecast describes AI-enabled quality control, predictive maintenance and autonomous material handling in manufacturing as increasingly dependent on real-time edge processing (IDC, 25 February 2026). Real time has a number attached, and for a language model the number a vendor quotes is usually tokens per second. The trouble starts when two vendors mean different things by it.

Three speeds under one name

llama.cpp's benchmark tool reports separate tests for prompt processing, text generation and the two combined, each as an average in tokens per second (llama-bench README). The README adds that results vary with batch size, model size, hardware and how many layers sit on the GPU.

Prompt processing is reading what you sent. Generation is writing the answer, one token after another. They run at different speeds, and they matter for different jobs. A shift summary built from a long maintenance log is dominated by reading. A short operator reply is dominated by generation. A speed claim that does not say which one it is leaves you comparing a reading rate with a writing rate.

vLLM's benchmark documentation draws more distinctions. It defines time to first token as the time from sending a request to receiving its first streamed output, inter-token latency as the time between consecutive streamed outputs, and throughput in requests and tokens per second across a run (vLLM benchmark CLI). The same page notes that with speculative decoding one streamed output can hold several tokens, so inter-token latency and per-token time stop agreeing.

A rate needs a stream

Look at how those definitions are phrased: "first streamed output," "consecutive streamed outputs." A rate is a count over a span of time, and the span needs at least two arrival times. llama.cpp's server sends tokens as they are predicted when a request sets stream, using server-sent events, and otherwise returns the finished reply (llama.cpp server README).

A reply that arrives whole gives one timestamp. You learn how long the caller waited and nothing about how that wait divides.

An illustration, with round numbers chosen only for the arithmetic: a 200-token answer comes back after 10 seconds. Dividing gives 20 tokens per second. If 6 of those seconds went to reading a long prompt, generation ran at 50 tokens per second, and the claim of 20 describes neither stage. With streaming you see the first token arrive at second 6 and the remaining 199 tokens spread over the other 4 seconds. Without it you see 10 seconds.

This matters beyond benchmarks. Any system that picks a model or a machine by its measured speed is picking on whatever the measurement can see. Integration layers sit between the client and the engine, and one that waits for the full reply and then passes it on removes the stream before anyone times it. Streaming has to survive every hop to be measurable at the end.

Why a quoted speed rarely matches a plant workload

A figure can come from a short prompt, one request and a warm cache, which flatters generation speed and hides reading time. A plant deployment sends long prompts, several requests at a time and a cold start after a reboot. The memory limits that decide how many requests fit are covered in the KV cache arithmetic note, and the engine question is in llama.cpp or vLLM.

What to ask when a vendor quotes a speed

  1. Is the figure prompt processing, generation or both combined?
  2. How long was the prompt, and how long the answer?
  3. How many requests ran at once?
  4. Was the reply streamed, and can I see time to first token as its own number?
  5. Which model file and quantization, on which exact module?
  6. Is the figure an average, and over how many runs?

If the answer to the fourth question is no, the first and third have no way to be checked.

llama.cpp's llama-bench README, the llama.cpp server README, vLLM's benchmark documentation and IDC's edge spending release (25 February 2026), all read on 6 October 2026. The worked example is illustrative arithmetic, not a measurement.

Related notes

Related paper: Combining agentic AI with distributed orchestration at the industrial edge