Will a Local LLM Fit? Check the Model Metadata
Vendors state capacity as parameters or gigabytes. A fit check needs three numbers from the model's own metadata, and a guess from parameter count can miss by a factor of four.
By Harinderpal Hanspal on September 2026. Updated October 2026
Vendors state what fits as a parameter count or a memory size, with no context length. A fit check that means anything needs the layer count, the key-value head count and the head dimension, which a GGUF file carries in its header. Without them the honest answer is unknown.
NVIDIA states that DGX Spark has 128 GB of unified memory and can run inference on models of up to 200 billion parameters (NVIDIA DGX Spark). OpenAI's repository says gpt-oss-120b runs on a single 80 GB GPU and gpt-oss-20b within 16 GB (gpt-oss README). Both are capacity claims. Neither names a context length or a number of simultaneous sessions, and the DGX Spark pages we read did not state the weight precision behind the 200 billion either.
A buyer deciding whether an edge box can host a model needs a sharper answer than that.
What a fit claim leaves out
The weights are one part of the memory bill. The other is the KV cache, the per-token memory the model keeps so it does not recompute earlier text. The KV cache arithmetic note works it through for Llama 3.1 8B: the cache grows with context length and with each simultaneous session, and at long context it can outweigh the quantized weights.
So "fits in 16 GB" is really a question with three unknowns: which context length, how many sessions, and what else shares the memory. A claim that leaves them out may be true at 4,000 tokens and false at 64,000.
The three numbers a check needs
Sizing the cache takes the number of layers, the number of key-value heads and the head dimension. Many current models share key-value heads across groups of attention heads, called grouped-query attention, and the count differs by model. Llama 3.1 8B has 8 key-value heads against 32 attention heads, so a check that assumed one per attention head would overstate its cache fourfold, as that note shows.
The practical consequence: a parameter count or a model family name cannot stand in for the architecture. Two models of the same size can need very different cache memory, and a guess errs in whichever direction the guesser did not think of. The cost of an error is concrete. An operator either downloads several gigabytes and finds the box cannot hold the context, or passes over a model that would have run.
Where the numbers already are
The GGUF format used by llama.cpp stores them. Its specification places the metadata key-value pairs in the file header and defines architecture keys such as block_count (the number of attention and feed-forward blocks) and attention.head_count_kv (the heads per group in grouped-query attention) (GGUF specification). The same spec says an architecture's listed keys are required for that architecture, though not every key applies to every architecture.
Since the header comes first, a tool can read the architecture without loading the weights. Whether a given vendor's tool does so is a question to ask, because a model catalog that lists only a name and a file size cannot support a context-aware answer.
What a check can say without the header is smaller but real: the weights alone need at least the file size, plus runtime overhead. The llama.cpp quantize README lists Llama 3.1 8B at 4.58 GiB for Q4_K_M and 14.96 GiB for F16 (llama.cpp quantize README). That floor is the same at every context length, which is why it can rule a model out but never rule it in.
Engines add their own levers. vLLM's documentation advises limiting max_model_len and max_num_seqs to reduce memory, and notes that quantized models take less (vLLM, conserving memory). A fit verdict is therefore tied to the settings it assumed.
Questions to ask when a vendor says it fits
- At what context length, and with how many sessions at once?
- Which quantization, and what is the file size?
- Where did the cache figure come from: the model's layers, KV heads and head dimension, or an estimate from parameter count?
- What does the tool answer for a model whose architecture it cannot read?
- How much memory is left after the operating system and the other models on the board?
An answer of "unknown" to the fourth question is a good sign.
NVIDIA's DGX Spark and Jetson product pages and the gpt-oss README, the GGUF specification, vLLM's memory documentation and the llama.cpp quantize README, all read on 5 and 6 October 2026. No figure here is a measurement on hardware.
Related notes
- How much memory an edge AI box needs: the KV cache arithmetic
- The cost of running AI on your own hardware starts with updates you must reach
Related paper: The Future of Industrial Operations: What's Needed to Build a Turnkey Autonomous Edge Platform for the Software-Defined Industry