Edge AI Memory: KV Cache and Context Length
Worked from a published model configuration: 128 KiB of KV cache per token, 16 GiB at full context, plus the weights at each quantization, and the questions to ask a vendor who says a model fits.
By Harinderpal Hanspal on October 2026. Updated October 2026
Whether a language model fits an edge box depends on two sums: the weights, set by quantization, and the KV cache, set by context length and simultaneous sessions. For Llama 3.1 8B the cache is 128 KiB per token at 16-bit, 16 GiB at full context.
Every number below is arithmetic from published figures. None is a measurement.
Why memory is the first limit on an edge box
IDC put worldwide spending on edge solutions at $265 billion in 2025 and $450 billion by 2029 (IDC, 25 February 2026). It names AI-enabled quality control and predictive maintenance in manufacturing as depending increasingly on real-time processing at the edge. Those workloads run on modules that ship with a fixed amount of memory. NVIDIA lists 8 GB for the Jetson Orin Nano Super and 128 GB for Jetson AGX Thor (Orin Nano Super, Thor). A cabinet box cannot borrow memory from a neighbor, so the sizing has to be right before it is bought.
The KV cache, one token at a time
The published configuration of Llama 3.1 8B gives 32 layers, 32 attention heads, 8 key-value heads, a hidden size of 4096 and a head dimension of 128. It is stored in bfloat16, 2 bytes per value, and the maximum position is 131,072.
For each token, each layer keeps one key vector and one value vector per KV head:
2 (K and V) x 32 layers x 8 KV heads x 128 dimensions x 2 bytes = 131,072 bytes = 128 KiB per token.
Scale by context length:
| Context | KV cache at 16-bit |
|---|---|
| 8,192 tokens | 1 GiB |
| 32,768 tokens | 4 GiB |
| 131,072 tokens | 16 GiB |
The 8 KV heads matter. With one KV head per attention head, the model would have 32, and the cache would be 4 times larger: 512 KiB per token, 64 GiB at full context. A guess from parameter count can miss by that factor.
Add the weights
llama.cpp's quantize README lists the Llama 3.1 8B file sizes: F16 at 14.96 GiB, Q8_0 at 7.95, Q6_K at 6.14, Q5_K_M at 5.33, Q4_K_M at 4.58 and Q2_K_S at 2.78.
| Weights | Plus 8K cache | Plus 32K cache | Plus 128K cache |
|---|---|---|---|
| F16, 14.96 GiB | 15.96 | 18.96 | 30.96 |
| Q4_K_M, 4.58 GiB | 5.58 | 8.58 | 20.58 |
All in GiB. The totals leave out runtime overhead, activation buffers and everything else on the board. An 8 GB module is about 7.45 GiB, so the Q4_K_M model at 8K context uses three quarters of it before the operating system and the camera pipeline take their share, and at 32K context the sum already exceeds it.
Quantizing the weights from 14.96 to 4.58 GiB saves 10.38 GiB. Going from 8K to 128K of context at 16-bit spends 15 GiB. At long context the cache, not the model file, decides whether it fits.
Concurrency multiplies the cache
Each active sequence holds its own cache. Four simultaneous requests at 32,768 tokens each need 4 x 4 GiB = 16 GiB of KV cache on top of one copy of the weights, unless the engine shares a common prefix between sequences. A plant assistant that serves four operators, or an agent that fans out to sub-agents, lands in this regime.
vLLM's documentation points the same way. It advises limiting max_model_len and max_num_seqs when memory is short and notes that quantized models take less memory (vLLM, memory conservation). Those two limits are the context length and the concurrency in the sums above.
What to ask before you buy
A vendor who says a model fits a given box should answer five questions with numbers:
- Which quantization, and which file size?
- What context length was assumed, and what is the longest your use needs?
- How many sessions at once, and does the engine share a common prefix between them?
- How much memory is left after the operating system, the application and the other models on the board?
- Were these figures measured on this exact module, or calculated?
A calculation is a fine answer if it shows its inputs. An unmeasured claim with no inputs is the one to push on.
Arithmetic derived from the published Llama 3.1 8B config.json (an unsloth mirror of the original), file sizes from llama.cpp's quantize README, vLLM's memory documentation and NVIDIA's Jetson specifications, all read on 5 and 6 October 2026. No figure here is a measurement on hardware.
Related notes
- Will this model fit my edge box? Why a parameter count cannot say
- The cost of running AI on your own hardware starts with updates you must reach
Related paper: The Future of Industrial Operations: What's Needed to Build a Turnkey Autonomous Edge Platform for the Software-Defined Industry