Valid JSON From Local LLMs: llama.cpp and Ollama
Berkeley's function-calling work shows output format alone moves small-model accuracy. llama.cpp and Ollama can constrain output to a schema, which fixes syntax and nothing else.
By Harinderpal Hanspal on August 2026. Updated October 2026
Small local models are sensitive to output format: Berkeley's function-calling study found 8B models degrade sharply when asked for XML-style tags. llama.cpp grammars and Ollama's format field constrain output to a JSON schema, which guarantees syntax, not that the chosen tool or values are right.
When Berkeley's Gorilla group varied only how models were asked to return a function call, accuracy moved. The study covered 39 models, 26 prompt variations and 200 single-turn test entries, and found performance "higher when the model is prompted to return function calls in Python or JSON format" than in either XML format (Gorilla, BFCL V4 prompt variation, 17 July 2025). Some small models, among them Llama-3.1-8B-Instruct and BitAgent-8B, showed "significant performance drops" when tags were required.
That result concerns tag formatting, not overall accuracy. It is still the right warning for anyone putting an 8B model at a plant: how you ask for the output is part of the system.
Why free-text JSON fails and what a grammar changes
Ask a model for JSON in the prompt and you get JSON most of the time. Most of the time is the problem. A downstream parser meets a stray sentence, a missing brace or a trailing comma, and a control loop has to decide what to do.
Constrained decoding moves the check inside generation. llama.cpp's grammar format, GBNF, is "a format for defining formal grammars to constrain model outputs", and its README says it can force valid JSON. A JSON schema is converted to a grammar. The server takes it as a json_schema body field or the response_format field on /chat/completions, and the command-line tools take --json (llama.cpp, grammars README). The model cannot emit a token that would break the schema.
Two details in that README matter in practice. The schema is used only to constrain output and "is not injected into the prompt", so the model still needs the field meanings described in words. And the README lists gaps: additionalProperties defaults to false, and uniqueItems, contains and conditionals are unsupported. It also warns that grammars "currently have performance gotchas", so measure speed on the device you are buying.
The Ollama route
Ollama takes a JSON schema in the format field of a request. Its documentation suggests also passing the schema as a string in the prompt "to ground the model's response", defining schemas with Pydantic or Zod "so they can be reused for validation", and lowering temperature, for example to 0. It works through the OpenAI-compatible API with response_format, and Ollama's cloud service does not support structured outputs (Ollama docs).
Read those tips closely. A vendor that tells you to validate has told you the guarantee has an edge.
What a grammar does not check
A grammar proves the output has the right shape. It does not prove the call is the right one. A valid {"action": "close_valve", "id": 7} can name the wrong valve. That limit is our inference, not the README's, and it is why CISA's guidance on AI in operational technology talks about error-checking that keeps an agent's outputs "within the expected bounds" (CISA and partners, 3 December 2025). The shape check belongs in the engine. The meaning check belongs in software that knows the plant: allowed values, ranges, which asset the caller may touch.
Schema design carries part of that. Enumerate the permitted actions rather than leaving a free string, bound numbers, and forbid extra fields.
A test to run on your own model and hardware
Take 100 requests representative of your use. Run each twice, unconstrained and with the schema. Count three things: outputs that fail to parse, outputs that parse but fail your semantic rules, and tokens per second. Repeat at the quantization you will ship, since a smaller file changes behavior. The result is a number for your model, which no benchmark page supplies.
Questions for a vendor supplying a local model with structured output:
- Is output constrained by the engine, or requested in the prompt and parsed afterward?
- What happens to a request the engine cannot constrain: an error, or a silent fall back to free text?
- Which schema features does the engine ignore or reject?
- What validates the values after the shape is guaranteed?
- Was speed measured with constraints on, on this hardware?
The wider argument for checks that fail loudly is in our paper on agentic AI at the industrial edge.
Drawn from the Gorilla (UC Berkeley) BFCL V4 prompt-variation study of 17 July 2025, the llama.cpp grammars README, Ollama's structured-outputs documentation and CISA's joint guidance of 3 December 2025, all read on 6 October 2026. No leaderboard score is cited.
Related notes
- llama.cpp or vLLM for AI agents on local hardware: concurrency decides
- Swapping the model under an agent: how tool-call formats differ
Related paper: Combining agentic AI with distributed orchestration at the industrial edge