llama.cpp vs vLLM for Agents: Concurrency

One person running an agent that fans out to sub-agents sends several requests at one model at once. What the llama.cpp and vLLM documentation say about that, and what to ask a vendor.

By Harinderpal Hanspal on October 2026. Updated October 2026

llama.cpp and vLLM suit different loads, and the deciding variable for an agent is how many requests hit one model at once. vLLM's authors say KV cache memory limits the batch size. No published figure here compares the two engines, so a vendor's pick needs a test on your hardware.

Which engine is faster depends on how many requests arrive at once, and no published benchmark gives the crossoverSketch of one parent agent fanning out to four requests that all hit one model. Below, an axis of concurrent requests shows two engine curves that may cross at an unknown point, marked with a question mark. A red note says no published benchmark gives the crossover. A side note says cache memory limits how many requests run together.parentagentonemodel4 to 8+ at oncecache memory limitshow many run together1manyrequests at oncecrossover?no published crossovertest on your ownhardware beforetrusting a pick
Which engine is faster depends on how many requests arrive at once, and no published benchmark gives the crossover

The paper that introduced vLLM starts from a memory problem. Its authors write that the KV cache for each request "is huge and grows and shrinks dynamically," and that when it is managed badly the waste limits the batch size (Kwon et al., arXiv:2309.06180, 12 September 2023). They report 2 to 4 times the throughput of the serving systems of that time at similar latency. Those systems were not llama.cpp, and the paper is three years old. What it establishes is the mechanism: how many requests a machine can serve together is set by cache memory, and the engine decides how well that memory is used.

For a plant or an OEM, the question arrives as a purchase decision. A box runs one model for an operator assistant, or for an agent that reads a fault, calls three tools in parallel and asks a sub-agent to summarize each result. Which engine sits underneath?

One chat session and a fan-out are different loads

A single user typing to a model sends one request, waits for the answer, and sends another. The engine has one sequence to run. An agent that fans out sends four, eight or more requests in the same second, and nearly all of them open with the same system prompt and the same tool definitions.

That shape raises two separate questions. Can the engine run the requests together in one batch? And can it avoid recomputing the shared opening for each of them?

What each project's documentation says

llama.cpp's server README lists parallel slots (-np, default auto), continuous batching (on by default) and prompt caching that re-uses the KV cache from an earlier request where possible (on by default) (llama.cpp server README). So the engine is not limited to one request at a time. The README does not say how throughput changes as slots fill, and each slot needs its own share of memory, which is the arithmetic in the KV cache arithmetic note.

vLLM's documentation describes automatic prefix caching: it keeps the KV blocks of processed requests and reuses them when a new request arrives with the same prefix, and the page calls this "almost a free lunch" and says it does not change model outputs (vLLM, Automatic Prefix Caching). For fan-out traffic with a shared opening, that is the feature to ask about, in whichever engine.

Speculative decoding is the other claim vendors lean on. vLLM's documentation frames it as a way to reduce inter-token latency at medium to low request rates, and points to simpler variants for peak traffic (vLLM, speculative decoding). Reading that page, the technique helps a lightly loaded box answer faster. It is not evidence of more throughput for a busy one, so the two speedups should not be added together.

Where a single-engine benchmark stops helping

A crossover point between the two engines is the number buyers want, and it depends on hardware, model, quantization and request shape. We found no primary-source benchmark that puts llama.cpp and vLLM on the same model and the same machine at several concurrency levels, so we quote no crossover.

Edge hardware adds a second break. vLLM's own memory guidance advises limiting the context length and the batch size when memory is short (vLLM, conserving memory). On an 8 GB module, those limits can leave little room for any batch at all, and the engine that serves fan-out well on a data-center GPU may have nothing to batch. The hardware and the engine have to be chosen together.

The test a buyer can ask for

vLLM's benchmark documentation defines the numbers worth requesting: time to first token, the time between consecutive streamed outputs, and throughput in requests and tokens per second (vLLM benchmark CLI). Ask for them at 1, 4, 8 and 16 simultaneous requests, on your hardware, with a prompt that opens like your real traffic.

  1. Which engine, which version, and which model and quantization were measured?
  2. At what concurrency does the quoted throughput hold, and what happens to time to first token as it rises?
  3. Do the test requests share a long opening prompt, as an agent's sub-agent calls do?
  4. Is prefix reuse switched on, and was it on in the test?
  5. How much memory does each additional simultaneous request cost on this module?
  6. If a second operator opens a session during a sub-agent run, what happens to the first?

A single tokens-per-second figure with no concurrency attached describes one request, at best.

vLLM paper abstract (arXiv:2309.06180, 12 September 2023), vLLM documentation on prefix caching, speculative decoding and benchmark metrics, and the llama.cpp server README, all read on 6 October 2026. No figure here is a measurement, and none compares the two engines.

Related notes

Related paper: Overcoming challenges in industrial edge computing: key solutions for scalable, secure operations