AI Data Classes: Which Requests Can Leave the Site

Classify the request, decide where each class may be processed, then make something other than a policy document stop the rest. Evidence from 2025 and 2026 telemetry, and a fill-in sheet.

By Harinderpal Hanspal on October 2026. Updated October 2026

A request may go to a hosted model only if its data class is cleared for that destination, and a control other than a policy document has to enforce it. Cyberhaven's 2026 report says 39.7% of data movements into AI tools involve sensitive data.

Each request carries a data class, and a gate decides where it may goSketch of a request tagged with a data class entering a gate. The gate sends it to a cleared destination or to a refused path. A fallback model sits below the gate with an arrow back into it, so it passes the same check, with the red pen on the refused path.requesttagged with aclassclass clearedfor this place?cleared placerefusedfallback modelsame check again
Each request carries a data class, and a gate decides where it may go

Cyberhaven's 2026 report says "39.7% of all data movements into AI tools involve sensitive data" (Cyberhaven, 5 February 2026). The unit is a data movement, which is not a prompt. The figure comes from the telemetry of Cyberhaven's enterprise customers through 2025, and the release does not state the sample size. The same release says 32.3% of ChatGPT usage runs through personal accounts, which a company cannot see into at all.

Netskope's January 2026 report reads similarly from another vendor's data. The average organization in its telemetry, October 2024 to October 2025, saw 223 incidents a month of users sending sensitive data to AI apps, about double the prior year (Netskope). The figure was read in a search extract of Netskope's page, not the full report. Harmonic Security found organizations had an average of 23 AI tools they did not previously know about (Harmonic Security, 31 July 2025).

Three vendors, three datasets, one direction. Sensitive data is already moving, mostly through tools nobody chose.

Four classes, four places

Most sites already have data classes for documents. The question here is the same one applied to a request: what is in it, and which destination is cleared for that? Below is a starting position a plant could adopt and then change. It is an example, not a standard.

Class Example content Where it may be processed
Public Published manuals, marketing copy Any hosted model
Internal Shift notes, maintenance tickets Hosted, in a named region, under a business contract
Confidential Process parameters, supplier pricing, customer contracts Your own data center or private cloud account
Restricted Controller programs, safety cases, personal data On a machine inside the site

The destination column is where plants disagree, and it should be. A pharmaceutical line and a packaging line will not draw the same boundary.

Location is a setting you can read

Residency is configurable and priced, which means it can be written into a class. Microsoft's Foundry documentation separates "Global" deployments, which may process prompts in any geography where the model is deployed, from "DataZone" deployments that stay within a specified zone (for an EU resource, any EU member state) (Microsoft Learn, page dated 18 May 2026). Anthropic lists US-only inference at 1.1 times the standard token rates (as of 6 October 2026, Anthropic pricing). Both are vendor statements about their own services.

One rule falls out of this. If you cannot tell where a destination processes data, treat it as the most exposed place on the list.

What actually stops a request from leaving

A policy PDF stops nobody who is in a hurry. Something has to refuse. Five controls do that work, and none of them comes from a standard we could cite as prescribing it.

The published guidance points the same way without prescribing it. OWASP's sensitive-information risk lists personal data, financial details, health records, confidential business data, credentials and legal documents, and recommends access controls, sanitization and redaction (OWASP LLM02, 2025). CISA and partner agencies list data classification among the focus areas of their May 2025 AI data security page (CISA); only the page's topic list was read, not the full guidance. None of these says to send one class to a local model and another to a hosted one. That step is an application of their principles.

The EU Data Act does not settle the question either. It governs access to and sharing of data from connected products, and it is not a rule that plant data must stay on site (European Commission).

Fill-in sheet: one row per data class

Copy this table and complete it with plant, OT, security and legal in the same room.

Class Real examples from your site Allowed destinations What refuses a wrong destination Who may grant an exception
Public
Internal
Confidential
Restricted

If the fourth column holds the words "staff training" or "policy", the row is not finished. The paper on governing agents in production argues for answering the data question in the layer every call must pass, below the agent.

Drawn from Cyberhaven's 2026 AI Adoption and Risk Report (5 February 2026), Netskope's Cloud and Threat Report 2026 (January 2026), Harmonic Security (31 July 2025), Microsoft Foundry and Anthropic documentation, OWASP LLM02 (2025), NIST AI 600-1 (July 2024) and CISA's AI Data Security page (22 May 2025), all read on 6 October 2026. The class scheme and the controls are ours.

Related notes

Related paper: Governing agents in production: what to ask before an agent acts