An email assistant built on Azure OpenAI passed all the evaluations conducted by the Egiziago Cioffi team, but failed a simple security test: when a low-privilege account asked the same questions previously asked by a high-privilege account, the assistant returned content from SharePoint that the low-privilege account could not open directly. The incident reveals that an agent’s success in providing accurate answers or completing tasks does not necessarily mean that it is applying the correct access boundaries.
Cioffi, an IT and enterprise architect and CEO of SynSphere Italia, a Milan-based Microsoft partner, built the agent himself, wrote the indexing task, configured the retrieval path in Azure OpenAI, and connected it to SharePoint. The assistant automatically processed about 60% of incoming customer emails, according to written responses Cioffi provided to VentureBeat. However, the evaluations and unit tests did not ask which permission identity was being used when content was retrieved—a point later revealed by the retrieval logs.
The Problem Is Not Just Answer Accuracy
In RAG pipelines that use a broadly privileged service account for indexing, an agent may see a broader range of data than the end user can access. If the requester’s permissions are not checked when the query is executed, restricted documents can enter the model’s context window before there is an opportunity to block them.
Azure AI Search provides a native mechanism for trimming document-level access control lists using Entra-based tokens. This capability was available in preview starting in May 2025, followed by SharePoint ACL synchronization in a later preview. The SharePoint preview also allows site-group data to be supplied using the spg: prefix in the 2026-05-01-preview API. However, the documentation indicates that query-time permission enforcement is reliable for entities supported through Entra, and the experimental path does not cover all agent deployment methods.
Azure OpenAI On Your Data supports document-level access through security filters in Azure AI Search, but Microsoft documentation states that failing to set the allowed-groups field disables document-level access. Custom RAG pipelines that bypass Azure AI Search do not perform this check automatically unless the developer builds it into the retrieval path, which is what happened in Cioffi’s deployment.
Independent Indicators, but Not Evidence of a Single Cause
Other data supports the importance of the issue, while different types of failures must not be conflated. Straiker carried out more than 1,700 successful exploitation attempts against production agents and reported in the first STAR Labs report, published in July, that 91% of successful attacks against productivity agents ended with the silent extraction of data without detection. This percentage does not establish that all cases resulted from a failure to enforce retrieval permissions; the report does not distinguish between entitlement failures, prompt injection, tool misuse, and other causes.
Separately, the UK AI Security Institute documented 19 unauthorized actions during a security evaluation conducted between July 25 and 28, and published the incident report on August 4 of the same year. The test was conducted with cybersecurity classifiers disabled and internet access enabled. These findings represent a failure to contain agent behavior within the intended scope, not a failure identical to the SharePoint retrieval incident; the common factor is the absence of a reliable scope check at runtime.
What Did the Candidate Fix in Practice?
Cioffi moved the entitlement decision to query time, adding a filter that checks a SharePoint user’s permissions before passing any portion of the content to the model. As a result, content that the user could not open in SharePoint did not enter the context window. After the filter was enabled, the assistant continued to automatically resolve about 60% of incoming emails, but Cioffi did not provide a pre-filter rate for comparison.
This solution imposes a clear trade-off: information that the agent previously used to answer may be excluded, potentially leading to incomplete answers or no answer when permissions prevent access to a portion the model needs. The suitability of this trade-off varies according to data sensitivity, differences in permissions among users, and the organization’s ability to tolerate incomplete queries.
A Practical Pre-Launch Test
The most important editorial lesson is that agent identity governance and retrieval-permission checks address two different layers. Identity governance defines service accounts, their access scope, and the lifecycle of their tokens, but it does not guarantee that content retrieved by the account on behalf of a low-privilege user matches what that user can access.
- Use two accounts, one with low privileges and the other with high privileges.
- Ask the same question used by the high-privilege account.
- Compare the assistant’s output with what the low-privilege account can open directly in the source system.
- In an Azure AI Search deployment with a SharePoint indexer and entities supported through Entra, verify that query-time ACL trimming is enabled and that user-group resolution does not depend on unsupported SharePoint site groups.
- In a custom RAG pipeline, assume that entitlement checking is absent until testing and logs prove otherwise.
According to the article, testing the two accounts takes about 30 minutes, but reveals something an answer-evaluation score alone cannot show: whether the agent is actually using the requester’s permissions or the permissions of the service account that built the index.