Nothing about Copilot creates an oversharing problem. What it does is find the one your file shares have quietly had for the last decade.
It is stated directly in the documentation: given the power and the speed with which AI surfaces content, generative AI amplifies both the problem and the risk of oversharing or leaking data. Purview supplies American businesses with the visibility and the controls, spanning Copilot, the enterprise AI applications you approved, and the consumer tools your staff open in a browser whether anybody sanctioned them or not.

- Three categoriesCopilot, enterprise AI apps, other AI apps
- EXTRACT rightWhat a label needs for Copilot to return data
- Prompts auditedCaptured in the unified audit log
- Browser DLPBlock pasting sensitive data into public AI sites
Seven things that decide whether your AI rollout is genuinely governed or simply deployed.
Three categories of AI application, not one
First, Copilot experiences and agents, which covers Microsoft 365 Copilot, Security Copilot, Copilot in Fabric and Copilot Studio. Second, enterprise AI apps, meaning non-Copilot applications connected through Entra registration, data connectors or Microsoft Foundry, with ChatGPT Enterprise and Anthropic Claude Enterprise both named. Third, other AI apps detected through browser activity, and that is where the consumer tools turn up.
Shadow AI is the third category, and the one that keeps people awake
Other AI apps are described as those detected through browser activity and categorized as generative AI within the Defender for Cloud Apps catalog, a group that uniquely includes applications built on third-party large language models. ChatGPT, Google Gemini, the consumer version of Microsoft Copilot and DeepSeek are all named. Most businesses currently have no visibility of any of it whatsoever.
Sensitivity labels, and the usage right that decides everything
Where a sensitivity label applies encryption, a user needs the EXTRACT usage right alongside VIEW before an AI app will return that data. That one detail is what allows a label to govern whether Copilot can surface a document at all, rather than governing only whether a person can open it. It is also frequently misconfigured, or missing entirely from label definitions that were written before anybody thought about AI.
A prerequisite most tenants have not met
Enabling sensitivity labels for SharePoint and OneDrive carries an explicit recommendation, and the consequence of skipping it is stated plainly. Without labels enabled for those services, the encrypted files Copilot and its agents can reach are limited to data in use from Office apps on Windows. That is a materially different protection posture from the one most businesses believe they already have.
Endpoint DLP that reaches the browser
Any Windows computer onboarded to Purview can carry endpoint data loss prevention policies that warn or outright block somebody sharing sensitive information with a third-party generative AI site through their browser. The published example leaves nothing to interpretation: a user is prevented from pasting credit card numbers into a public AI tool, or receives a warning they are able to override.
Prompts and responses are auditable, discoverable and retainable
The unified audit log captures prompts and responses, recording how and when a user interacted with the application, which Microsoft 365 service the activity happened in, references to any files accessed during that interaction, and whatever sensitivity label those files carried. All of it is stored in the user mailbox, which makes it searchable through eDiscovery and subject to your retention policies.
Risky AI usage as an insider risk signal
There is a risky AI usage policy template inside insider risk management, described as detecting risky usage including prompt injection attacks and attempts to access protected materials, with the resulting insights integrated into Microsoft Defender XDR. That gives AI misuse a route into exactly the same investigation flow as any other insider risk indicator you already handle.
Those permissions were always wrong. Copilot simply happens to be the first thing fast enough to notice.
The risk is framed precisely in the documentation, and that framing explains why AI readiness work turns out to be almost entirely data governance work.
- The wording runs like this: because of the power and the speed with which AI can proactively surface content, generative AI amplifies both the problem and the risk of oversharing or leaking data.
- The AI applications Purview supports rely on your existing controls to ensure data held in your tenant is never returned to somebody without access to it. That is genuinely reassuring, and it is not the problem. The problem is how many people already have access to things nobody ever intended them to have.
- A SharePoint site shared with the entire organization eight years ago carried little risk while finding anything inside it required knowing the site existed at all. It becomes a very different proposition once an assistant can summarize its contents in response to a question asked in plain English.
- The practical consequence is that your AI readiness project and your oversharing remediation project are one and the same. Businesses treating them as two separate things deploy Copilot, discover the problem during week two, and pause the rollout while everybody argues about whose problem it is.
Four things that keep an AI rollout moving instead of paused indefinitely.
We look before we plan
Three questions: which AI applications are genuinely in use, what is currently overshared, and what sensitive data is already appearing in prompts. Every one of those answers reshapes the work ahead. Build a rollout plan before knowing them and you have built a plan that gets revised during week two, usually in front of the same people who approved the original.
We check the label prerequisites nobody checks
Two in particular. First, whether sensitivity labels are enabled for SharePoint and OneDrive, because without that the encrypted files Copilot can reach are limited to data in use from Office apps on Windows. Second, whether your encrypting labels grant the EXTRACT usage right, since without it an AI app cannot return the data even to somebody entirely authorized to see it.
We warn before we block on consumer AI
Endpoint DLP is able to block somebody pasting sensitive information into a third-party generative AI site, and it can equally warn them with an override available. Beginning with warn produces data about who needs what and for what reason. Beginning with block produces exactly the same activity on a personal phone, where you have no visibility, no policy and no record of any of it.
We settle the retention and discovery questions early
Because prompts and responses sit in the user mailbox, they are discoverable through eDiscovery and fall under your retention policies. Deciding how long to keep them, and testing the eDiscovery query path before anybody urgently needs it, is a small piece of work that becomes distinctly awkward if you attempt it for the first time under legal pressure.
Four phases, and the visibility phase rewrites the plan every single time.
- 01Weeks 1 to 3
See what is actually happening
We establish which AI applications are genuinely in use across all three published categories, browser-detected consumer tools nobody sanctioned very much included. Then what data is actually overshared, and which sensitive information types are turning up inside prompts and responses. This phase is deliberately observational, and it reliably produces the one finding that reframes the entire project.
- An inventory spanning Copilot experiences, enterprise AI apps and the browser-detected ones
- Oversharing exposure identified against the data that matters
- Sensitive information types found in prompts and responses
- A realistic picture of shadow AI usage rather than an assumed one
- 02Weeks 4 to 8
Fix the data before enabling more AI
We enable sensitivity labels for SharePoint and OneDrive, because without them the encrypted files Copilot can reach are confined to data in use from Office apps on Windows. Label definitions get checked for the EXTRACT usage right wherever encryption is applied. Then comes the oversharing remediation itself, which accounts for the majority of the work in this phase.
- Sensitivity labels enabled for SharePoint and OneDrive
- EXTRACT usage right reviewed on every encrypting label
- Oversharing remediated on the highest-exposure sites first
- Labeling extended to the content that matters most
- 03Weeks 9 to 12
Put controls around consumer AI
Endpoint DLP policies go onto the onboarded Windows devices, warning on or blocking sensitive information being shared with a third-party generative AI site through a browser. We normally begin in warn mode with a business justification attached, because a hard block on day one simply drives people onto a personal device where you have no visibility whatsoever.
- Endpoint DLP policies scoped to the sensitive information types that genuinely matter
- Warn with justification before any hard block
- Policy tips written so that people understand them rather than route around them
- Approved alternatives communicated alongside the restriction
- 04Ongoing
Govern it like any other data
Prompts and responses are audited into the unified audit log, remain discoverable through eDiscovery because they live in the user mailbox, are retained or deleted according to your retention policies, and can be reviewed through communication compliance. Risky AI usage is detected as an insider risk signal, with the resulting insights integrated into Defender XDR.
- Retention decided for prompts and responses, not left undefined
- eDiscovery query path tested before it is needed
- Risky AI usage policy enabled where appropriate
- Compliance Manager AI regulatory templates assessed
Six US situations where AI data posture is the blocking issue.
A company that paused its Copilot rollout after the pilot
We hear this story more than any other. The pilot worked well, then somebody asked a question that surfaced a document they had no business seeing, and the whole program stopped. The remediation required is oversharing work rather than AI work, and framing it that way is what allows the rollout to resume from a defensible position rather than simply being delayed.
A regulated business that cannot allow client data anywhere near a public model
A financial firm under GLBA and the FTC Safeguards Rule, or any business carrying confidentiality obligations to its clients, needs rather more than a policy memo. Endpoint DLP on Windows devices onboarded to Purview will warn or block somebody sharing sensitive information with a third-party generative AI site through their browser. Put that together with visibility of which AI applications are genuinely in use and a policy statement becomes an enforced control with evidence sitting behind it.
A business asked to say what its staff have typed into AI tools
The question arrives through a customer security questionnaire, an insurance application, or from the board itself. Prompts and responses are captured in the unified audit log, recording how and when users interacted, which Microsoft 365 service the activity occurred in, references to whichever files were accessed, and any sensitivity label carried by those files. That constitutes a real answer rather than a quotation from a policy document, and it is available whether or not anybody has thought to look yet.
A healthcare organization worried about PHI reaching AI tools
Protected health information pasted into a consumer chatbot is a disclosure nobody authorized. Endpoint DLP policies keyed on the sensitive information types that match PHI can warn or block that action in the browser, and prompts and responses for sanctioned AI apps stay auditable, discoverable and subject to retention. Your compliance advisors own the HIPAA interpretation; we own the controls and the evidence.
A business worried about prompt injection and misuse
Insider risk management ships a risky AI usage policy template, described as detecting risky usage including prompt injection attacks and attempts to access protected materials, with the insights integrated into Defender XDR. That gives AI-specific misuse the same investigation path as every other insider risk signal, rather than requiring a separate process nobody has built.
A business preparing for AI regulation
Regulatory templates in Compliance Manager exist to assess, implement and strengthen the compliance requirements around generative AI applications, with monitoring AI interactions and preventing data loss in AI applications both named as examples. Which of those templates apply to your business, given the state-level AI laws now emerging and whichever sector you operate in, is something we work through against your actual obligations rather than against a general list.
How US businesses are governing AI use today.
| Feature | Governed AI posture | Copilot deployed, consumer AI banned on paper | No AI governance |
|---|---|---|---|
Copilot interactions audited | Yes | Available but unused | No |
Consumer AI usage visible | Yes | No | No |
Oversharing assessed before rollout | Yes | Rarely | No |
Sensitivity labels control AI access | Yes | Partly | No |
Pasting sensitive data into public AI blocked | Yes | No | No |
Prompts discoverable in eDiscovery | Yes | Untested | No |
Prompt retention decided | Yes | No | No |
Risky AI usage detected | Yes | No | No |
AI regulatory position assessed | Yes | No | No |
Answer to what did staff share with AI | Evidence | Assumption | None |
How each protection method behaves with AI applications.
Protection method
Sensitivity label with encryption
- Behavior with AI apps
- An AI app returns the data only where the user holds EXTRACT alongside VIEW
Protection method
Sensitivity labels not enabled for SharePoint and OneDrive
- Behavior with AI apps
- Copilot and its agents can reach encrypted files only as data in use from Office apps on Windows
Protection method
Azure Rights Management encryption without a label
- Behavior with AI apps
- Both VIEW and EXTRACT are still checked, though protection does not automatically carry across to new items
Protection method
S/MIME protected email
- Behavior with AI apps
- Never returned by Copilot, and Copilot itself is unavailable in Outlook while one is open
Protection method
Password-protected documents
- Behavior with AI apps
- Unreachable by an AI app unless the user already opened it in that same app, and the password does not carry over
Protection method
Customer Key or bring your own root key
- Behavior with AI apps
- Fully supported, with those items eligible for Copilot to return
Protection method
Endpoint DLP on onboarded Windows devices
- Behavior with AI apps
- Able to warn on or block sensitive information being shared with a third-party generative AI site through a browser
Protection method
Sensitive information types and trainable classifiers
- Behavior with AI apps
- Used to locate sensitive data inside prompts and responses, surfacing in the Purview reports and in activity explorer
Five steps, and every surprise lives in the first one.
- 1
Assess AI usage across all three categories
Three of them. Copilot experiences and agents. Enterprise AI applications connected through Entra registration, data connectors or Foundry. And other AI apps detected through browser activity and categorized as generative AI. That third group is where the uncomfortable findings live, and most businesses have never once looked at it.
- 2
Assess the data exposure that AI would amplify
We establish what is currently overshared, what carries a label, what carries none at all, and which sensitive information types are already appearing in prompts and responses. Since the risk is framed as AI amplifying an existing oversharing problem rather than creating a new one, the assessment gets written that way too, so that the remediation ends up scoped honestly.
- 3
Fix the protection prerequisites
Sensitivity labels get enabled for SharePoint and OneDrive, because without that the encrypted files Copilot and its agents can reach are limited to data in use from Office apps on Windows. Encrypting labels are checked for the EXTRACT usage right. And content protected by Rights Management without a label gets identified, since protection there does not automatically carry across to new items.
- 4
Put controls around unsanctioned AI
Endpoint data loss prevention goes onto the onboarded Windows devices, warning on or blocking sensitive information being shared with third-party generative AI sites through a browser. Warn with a business justification comes first. Block applies where no legitimate use exists. And an approved alternative is always communicated at the same moment as the restriction.
- 5
Govern the interactions as data
Retention policies get applied to prompts and responses, with any conflicts resolved through the principles of retention. The eDiscovery query path is tested rather than assumed. Communication compliance policies extend to AI interactions wherever message supervision applies to you. The risky AI usage insider risk template goes on. And the Compliance Manager AI regulatory templates are assessed against the obligations you actually carry.
What US businesses ask about AI data security posture.
Fifteen questions worth answering honestly.
Visibility
- Which AI apps are actually in use?All three published categories.
- Do you see browser-based consumer AI use?That is the third category.
- What sensitive data appears in prompts?Classifiers surface this.
- What is overshared today?The answer predates Copilot.
- Are AI activities visible in activity explorer?There is a dedicated tab.
Protection
- Are sensitivity labels enabled for SharePoint and OneDrive?Without it, coverage is limited.
- Do encrypting labels grant EXTRACT?Required for AI apps to return data.
- Is any content protected without a label?No automatic inheritance for new items.
- Are endpoints onboarded to Purview?Required for browser DLP.
- Is there an approved AI tool people can use?Blocking without an alternative fails.
Governance
- How long are prompts and responses retained?Retention policies apply to them.
- Can you find them in eDiscovery?They sit in the user mailbox.
- Are AI interactions in scope for message review?Communication compliance covers them.
- Is risky AI usage monitored?There is a policy template for it.
- Which AI regulations apply to you?Compliance Manager has templates.
The pages around this one.
Establish what your people are already typing into AI tools nobody ever approved.
That third application category exists for precisely this question, and hardly any business has ever looked at it. The assessment is short, and what it finds usually determines how the rest of your AI program gets sequenced.
Related Services
Explore more solutions that work great with this service
Microsoft Purview
Data governance and compliance solutions
Learn moreMicrosoft Purview Endpoint DLP
Endpoint data loss prevention for US organizations: device onboarding
Learn moreDLP Solutions
Data Loss Prevention implementation for US businesses via Microsoft
Learn moreMicrosoft Copilot
AI-powered productivity with Copilot
Learn moreCopilot for Security
AI-driven security operations
Learn moreIT Compliance
HIPAA, SOC 2, NIST, CMMC, CCPA readiness
Learn more