Open vs Closed Weight Models: Data Policy First
Weights decide which places a model can run. The data decides which of those places are allowed. Evidence on shadow AI, licences and the quality gap, and five questions to answer first.
By Harinderpal Hanspal on July 2026. Updated October 2026
Choosing between open-weight and closed-weight models is first a question about where a request's data may go. In IBM's 2025 study, one in five breached organizations reported a shadow-AI breach, and a hosted open model sends data out just as a hosted closed one does.
One in five organizations that suffered a breach in IBM's 2025 study reported that it involved shadow AI, meaning AI tools staff used without approval. The study covered 600 breached organizations worldwide between March 2024 and February 2025, and IBM reports that only 37% have policies to manage or detect shadow AI (IBM, 30 July 2025). Organizations with high shadow-AI use had breach costs $670,000 higher than those with low or none. That is a difference between two groups, not the price of one breach.
Harmonic Security looked at the prompt side. Across 1 million prompts and 20,000 uploaded files in more than 300 AI apps, April to June 2025, it flagged 4.37% of prompts and 22% of files as containing sensitive content (Harmonic Security, 31 July 2025). Both are vendor datasets drawn from their own customers, so read them as evidence the leak exists, not as a rate for your plant.
Why the weights debate hides the policy question
Teams usually argue open versus closed on quality and price. Epoch AI measured the quality side: since January 2026, the most capable open-weight models have "lagged frontier closed models by an average of four months", about 8 points on its aggregate capabilities index (Epoch AI, 29 May 2026). The index says nothing about any single task. A maintenance-summary job might see no gap, and a hard reasoning job might see a wide one.
Now put a request in front of that choice. A shift handover note that names a customer, a recipe parameter or a supplier price has a policy answer before it has a quality answer. If the policy says it may not leave the site, the best closed model on the market is off the list, and so is a hosted copy of the best open one.
Open weights say nothing about where the model runs
The weights are a file. Where that file runs is a separate decision. An open-weight model rented from a hosting provider processes your request on the provider's machines, which is the same exposure as a closed model's API. An open-weight model on your own hardware keeps the request on site. A closed model cannot be served from your own hardware at all, so weights decide which places exist and data decides which of them are allowed.
Hosted does not automatically mean leaked, and the documentation shows why this is a configuration question. Microsoft says prompts and completions in Foundry are not available to OpenAI or other model providers and are not used to train foundation models without permission, but a "Global" deployment may process requests in any geography where the model is deployed, while a "DataZone" deployment keeps processing inside the named zone (Microsoft Learn, page dated 18 May 2026). Amazon says model providers have no access to Bedrock logs or customer prompts and completions (AWS documentation). Anthropic prices US-only inference at 1.1 times the standard token rate (as of 6 October 2026, Anthropic pricing). Those are vendors' assurances about their own services, not audits, and they are priced and optional.
Licences are a second policy, and they differ by model
Choosing open weights adds a legal check. Meta's Llama 4 licence requires a company with more than 700 million monthly active users to request a license, and requires "Built with Llama" to be displayed (Llama 4 licence). Its use policy withholds the multimodal-model rights from individuals domiciled in, or companies with a principal place of business in, the European Union (Meta use policy). The policy says it does not restrict end users of products that incorporate those models. The Open Source Initiative has said the Llama 3.x licences are not open source (OSI, 18 February 2025).
Other families are plainer, as of 6 October 2026: gpt-oss, Qwen3-235B and Mistral Large 3 carry Apache 2.0 on their model cards, and DeepSeek-R1 is MIT, though its Llama-based distilled variants inherit Llama terms (DeepSeek-R1 card). A shortlist built on "open" without reading the card can include a model your legal team cannot approve.
What the guidance says, and what it does not
No primary source we found says "send this class of data to a local model and that class to a hosted one." What exists is principle. OWASP's LLM02 list names personal information, financial details, health records, confidential business data, security credentials and legal documents as sensitive, and lists redaction, least-privilege access and sanitization as mitigations (OWASP, 2025). It does not recommend local models by name. NIST's generative AI profile warns that models "may leak, generate, or correctly infer sensitive information about individuals" and suggests documenting reliance on third parties and identifying fallbacks (NIST AI 600-1). Turning that into a model choice is your policy work.
The same pressure shows in adoption data. In a McKinsey and Mozilla Foundation survey of more than 700 technology leaders, security and compliance was the most cited barrier to open-source AI, at 56% (McKinsey). It is a self-selected sample of technologists, and the figure comes from a search extract of McKinsey's page.
Five questions to answer before a model is shortlisted
- Which kinds of data may never leave a site you control, and who signed that list?
- For each remaining kind, is a named processing region required, and who pays the premium for it?
- Does every model on the shortlist have a licence your legal team has read for this use, including any distilled variant?
- If the shortlist's best model is hosted, which request types stay behind and run on something you own?
- When a model is retired or swapped, who re-checks the data rule, not only the output quality?
The paper on governing agents in production argues for answering the data question in the layer every call must pass, below the agent.
Drawn from IBM's Cost of a Data Breach Report 2025 (30 July 2025), Harmonic Security (31 July 2025), Epoch AI (29 May 2026), the Llama 4 licence and use policy, model cards on Hugging Face and GitHub, Microsoft Foundry and Amazon Bedrock data documentation, OWASP LLM02 (2025) and NIST AI 600-1 (July 2024), all read on 6 October 2026. The framing is ours.
Related notes
- AI agent cost tracking: a spend cap cannot say which agent earned its bill
- The cost of running AI on your own hardware starts with updates you must reach
Related insights
Related paper: Governing agents in production: what to ask before an agent acts