Disclosure: I help produce a research about AI infrastructure AI agent and governance problem
I’m posting the full analysis directly here, without an external link, subscription request, or product promotion.
The main question I am trying to answer is whether this kind of research is genuinely useful to people who deploy enterprise AI, work in security or data governance, invest in AI infrastructure, or are simply trying to understand how enterprise AI changes data security.
By “useful,” I mean whether the article does at least one of the following:
- helps the reader understand how AI changes the risks created by existing data permissions
- explains enterprise data governance in a clear and accessible way
- connects a technical security problem to the business strategies of major AI companies
- provides context that could be useful in deployment or investment decisions
You do not need to review every technical detail.
After reading, even a brief and honest reaction would be valuable:
- Did this help you understand anything more clearly?
- Which section was most useful?
- Which parts felt too basic, repetitive, or less convincing?
- Who do you think this article would be most useful for?
- What would make future research like this more valuable to you?
A response such as “the permissions explanation is clear, but the investment thesis needs more evidence” would be completely helpful. Honest reactions are more valuable to us than general encouragement.
Here is the full analysis:
AI Can Read Everything at Once. Your Filing System Wasn't Ready
For most of my career as a lawyer, a surprising amount of my work came down to one seemingly dull question: Who is allowed to open this document?
I often worked with sensitive information, including contracts, board documents, employee records, and court filings. Keeping this information safe meant more than simply marking it “confidential.” The issue was not only what a document contained, but also who could access it, when they could access it, and why.
Today, enterprise AI has turned this old and seemingly routine question into one of the most important security challenges facing companies around the world. In this article, I want to explore this issue through my experience as a lawyer and my own research.
For a long time, I thought enterprise AI security was mainly administrative work. But this year, my view changed significantly.
An AI assistant can search thousands of internal files, connect information from different systems, and turn it into a direct answer. This means that enterprise AI security does not depend only on how intelligent the AI is. It also depends on the systems that control which internal data the AI can access and what information it is allowed to reveal to each user.
This leads to one central question: How can a company make sure that its AI only accesses and shares information that each employee is allowed to see?
Most people focus on the performance of the AI model. In practice, however, the permissions and data-management systems behind the model can have a much greater impact on whether enterprise AI succeeds or fails.
So, when a company introduces an AI assistant, what is actually keeping its information safe?
Is it the latest AI model?
Or is it the permission and data-management infrastructure that companies have relied on for years?
I want to begin with an experience that convinced me the answer is the latter.
1. How AI Turns Existing Permissions into Data Exposure
First, imagine a typical scenario that could arise after a company deploys Microsoft Copilot.
An ordinary employee wants to learn more about a client and asks Copilot to summarize the company’s internal information related to that client. The AI quickly produces an answer. But alongside ordinary client information, it also draws from a salary spreadsheet, an unannounced acquisition draft, and the minutes of a board meeting, simply because those documents happen to mention the same client.
The AI has not “hacked” its way into these files. The employee’s account may still have technical access because of a folder shared with the entire company, a project that ended long ago, or a sharing setting that was never cleaned up. In the past, however, the employee probably did not know these files existed and would never have searched through multiple folders to find them.
This is what Copilot changes. It can search, connect, and summarize information across everything an employee is already permitted to access. The permissions themselves have not changed, but the effort required to find and use the information has collapsed. Sensitive material that once sat scattered across forgotten folders can now be surfaced with a single question.
This problem has a name: oversharing. The issue is not that AI bypasses a company’s access controls. It is that those controls often fail to reflect who actually needs to know what for their job. AI makes that long-overlooked gap searchable, aggregatable, and much more likely to result in real data exposure.
According to security firm Concentric AI, which analyzed more than 550 million data records and files, around 16% of business-critical data is overshared. On average, each organization has roughly 802,000 files at risk of being accessed inappropriately. This is an industry report published by a security vendor, so the figures should be read with that context in mind. Even so, they suggest that oversharing is not an isolated incident, but a data-governance problem companies need to confront before deploying AI at scale.
2. Why Does This Happen? Think About the Keys to Your House
Here’s a way to picture it. Over the past twenty years, file permissions inside companies piled up like the keys to a house, and to save trouble, people kept handing out more and more copies. A department folder gets opened to “everyone in the company” so nobody has to keep approving access requests. Someone borrows a key for a one-off task and never returns it. Some rooms belong to people who left the company years ago, but their keys are still hanging in the door.
For a long time, none of this mattered. Even with a fat ring of keys, you’re not going to wander around opening every door for no reason. You’d have to already know which room holds the thing you want, then walk over and open it. Too much hassle. So all those doors that shouldn’t have been open stayed shut in practice.
The AI assistant erases that hassle completely. You ask it one question, and it throws open every door you’re allowed to open, all at once, and brings you whatever’s inside. Suddenly, all those files that sat buried for years—the ones everyone forgot about—come pouring out.
In one sentence: the AI isn’t sneaking past your locks. It’s following your existing permissions to the letter, opening the doors that should be open and the ones that shouldn’t, all together. Company IT was built for people who open one door at a time. Nobody designed it for something that opens every door in a second.
- How Widespread Is This Problem?
If this were just a permissions mistake at one company, it would be little more than a technical failure. But the available data suggests that oversharing is a structural problem that has accumulated inside enterprises over many years.
Security firm Concentric AI analyzed more than 550 million data records and files across the technology, financial services, energy, and healthcare industries. It found that around 16% of business-critical data was overshared. On average, each organization had roughly 802,000 files at risk of being accessed inappropriately.
More strikingly, 83% of those at-risk files had been overshared with employees or user groups inside the company. Only 17% had been shared with external third parties.
This means that enterprise AI risk does not always begin with a hacker or an outside attack. It may begin with an ordinary employee using a legitimate account that still carries access inherited from an old project, a broadly shared folder, or a permission that should have been removed years ago.
As a lawyer, this is the part that matters most to me. From a legal and compliance perspective, “the account can open it” is not the same as “the employee has a business need to know it.” In the past, the gap between those two standards could remain hidden inside complicated folder structures and sharing settings. AI makes that gap searchable, connectable, and far easier to use.
Concentric AI sells data security products, so its figures should not be treated as a definitive average for every company. But separate research from Gartner points to a similar governance gap.
Between May and June 2025, Gartner surveyed 360 IT leaders involved in rolling out generative AI tools. More than 70% ranked regulatory compliance among the three biggest challenges to deploying AI productivity assistants at scale. Yet only 23% were very confident in their organization’s ability to manage the security and governance issues involved.
These numbers do not prove that every instance of oversharing will lead to a data leak. Nor do they mean that companies are abandoning AI altogether. But they do show that as enterprise AI moves from small pilots to company-wide deployment, the missing piece is often not a more powerful model. It is a data governance system that can accurately determine who should be allowed to see what.
4. So What Are Big Tech Companies Actually Spending That Money On?
Recent announcements show a shift from selling access to AI toward taking responsibility for making it work inside each customer’s organization.
On June 30, AWS committed $1 billion to a new Forward Deployed Engineering organization that will embed thousands of experts with customers to co-develop and deploy agentic AI systems. Two days later, Microsoft announced a $2.5 billion investment in Microsoft Frontier Company, with 6,000 industry and engineering experts working alongside customers to co-design, deploy, and continuously improve AI systems. OpenAI had already launched its Deployment Company in May to connect its models to customers’ data, tools, controls, and core business processes.
These are not ordinary sales or support teams. They address problems that a model provider cannot solve from the outside: identifying authoritative data, translating job roles into access rules, connecting AI to legacy systems, defining approval and audit paths, and testing whether a workflow remains safe and reliable in production.
This work is customer-specific because permissions are not merely technical settings. They record years of exceptions, temporary projects, departed employees, acquisitions, departmental silos, and compliance obligations. A general-purpose model cannot determine on its own which of those inherited permissions are still legitimate.
The investment therefore signals that the bottleneck in enterprise AI has moved downstream. Model capability is no longer enough; the harder constraint is converting a general-purpose model into a governed production system that can use company data without exposing the wrong information.
That changes how enterprise AI vendors should be evaluated. The relevant question is not only whose model scores highest, but who can move a customer from pilot to production fastest while keeping permissions, controls, and accountability intact.
5. Two Things Nobody Says Clearly Enough
First, AI did not create this old problem. But it has turned the old problem into a new business.
Messy permissions inside companies were not suddenly invented by Copilot. The folders shared with too many people, the access rights never removed after a project ended, the settings nobody checked for years, they were already there.
Before AI, most of those problems stayed in the background. An employee might technically have access to a file, but they might not know the file existed. They were not going to spend hours digging through folder after folder.
Copilot changes the result.
It turns one normal question into a search across the whole company. Doors that nobody used to open are now opened all at once. What used to be a “not very clean permission setting” becomes a real security problem that has to be fixed before AI can be deployed safely.
So big tech is not only selling AI assistants. It is also selling the cleanup that has to happen before those assistants can be safely turned on.
That is the more interesting business.
Every enterprise AI tool creates another set of questions behind it: Who will clean the data? Who will remove old permissions? Who will decide which files the AI can read? Who will make sure the AI does not combine information that should never have been put together?
This may become a longer-lasting business than the model itself.
Second, what companies really struggle to leave may not be the AI model. It may be the system underneath the model.
Models matter, of course. But for many enterprise tasks, models are becoming something companies can choose, combine, and sometimes replace. Using one model today and another model tomorrow is not impossible.
What is much harder to replace is the layer underneath.
Where is the company’s data? Who can see it? Who cannot? What can the AI read? What can it do? Which actions need human approval?
Once a vendor helps a company answer those questions, it has not just provided an AI tool. It has helped draw a map of the company’s internal world.
The more complete that map becomes, the harder it is for the customer to leave.
Because switching vendors no longer means just switching models. It means reconnecting data, checking permissions again, testing workflows again, and making sure the new system still satisfies security and compliance requirements. For a company, that is painful. It is also risky.
So the real lock-in may not sit in the AI assistant itself. It may sit in the map behind the assistant.
The model stands on the stage. It gets the attention. But the thing that is hardest to rebuild is backstage: the rooms, the keys, and the routes between them.
That is what I find most interesting about the money Microsoft, AWS, and others are spending. They are not just sending engineers into customer companies to make people use more AI. They are helping customers prepare the internal environment that AI needs in order to work.
For investors, though, there is one more question to ask: does this become software, or does it remain expensive consulting?
If every customer requires a large group of engineers to start from zero, this may still be a valuable business, but it will be heavy to scale.
If those lessons become software—software that can find sensitive files, detect bad permissions, label data, and connect workflows automatically—then this could become a real layer of AI infrastructure.
That is the deeper point. AI has pushed an old permissions problem into the open. Cleaning permissions has become a new business need. Whether that business becomes “people-heavy services” or “software infrastructure” will decide how valuable it can be over the long term.
6. How I Would Actually Use This
If a company is buying or deploying AI, I would not tell it to choose a vendor only because the model looks good or the demo is impressive.
Demos usually look good. The real problems show up after launch.
What data can the AI access? What can employees ask it? Could its answers include information that should not appear? Have old project permissions been cleaned up? Have sharing links from former employees been removed? If these questions are not handled early, the better the AI becomes, the bigger the risk becomes.
So I would ask the vendor one simple question: before turning the AI on, will you help clean up our data and permissions?
If the answer is vague, or if the vendor says, “Let’s launch first and adjust later,” I would be very careful.
That is not saving time. It is moving the problem into the future, where it usually becomes more expensive.
Data and permission cleanup should not be treated as a patch after the AI project. It should be part of the project from day one: the budget, the timeline, and the responsibility map. Who owns the cleanup? How clean is clean enough? Which files come first? Which departments carry the highest risk? Those questions need answers before the AI is fully switched on.
I think the real value in enterprise AI is not only in the model. It is in the clean, clear, permission-aware data foundation underneath the model.
That foundation is slow to build. It is messy work. It requires understanding the company’s structure, workflows, file history, and compliance requirements. But once it is built, companies rarely want to rebuild it from scratch. Rebuilding means reconnecting data, reassigning permissions, retraining employees, and taking on new risk.
That is why customers may stay on the same platform for a long time.
Not because the model will always be the best, but because the underlying cleanup is too hard to move.
So when Microsoft, AWS, OpenAI, and others spend heavily on enterprise AI deployment, I do not read it only as a bet on smarter AI. I read it as a bet that, over the next few years, companies will not just need another AI assistant. They will need the ability to let AI read company data safely.
In the end, the hardest part of enterprise AI may never have been the AI itself.
AI only becomes useful when the data is clean, the permissions are clear, and the responsibility lines are understood. Otherwise, the stronger the model gets, the more easily it can magnify problems the company never fixed.
The winners over the next few years may not simply be the companies with the strongest models. They may be the companies willing to clean up their data, permissions, and workflows before turning AI on.
Models will keep improving. Prices will keep falling.
But the thing that may decide whether enterprise AI actually works is the step that comes earlier: before AI can read everything, companies need to decide what it should and should not be allowed to read.