Cybersecurity

Google: Indirect Prompt Injection Attacks Are Emerging on the Web but Remain Limited in Sophistication

A survey conducted by Google’s threat intelligence teams of public web content shows that various actors have begun experimenting with indirect prompt injection to target artificial intelligence systems, with the malicious category increasing by 32% between November 2025 and February 2026. Despite the limited sophistication of the observed attacks, Google expects their volume and complexity to grow as artificial intelligence systems become more capable and the cost of carrying out attacks declines.

2026-04-23
6 min read
12 views
فريق تحرير certi.news
Google: Indirect Prompt Injection Attacks Are Emerging on the Web but Remain Limited in Sophistication

An analysis conducted by Google’s threat intelligence teams concluded that indirect prompt injection has begun appearing in public web content as an experimental activity by attackers and website owners. However, most cases detected so far have been limited in sophistication and have not shown widespread reliance on advanced techniques. Nevertheless, Google recorded a 32% increase in detections in the malicious category between November 2025 and February 2026, a trend it considers an indicator of growing interest in this type of attack.

Google published this analysis on its security blog, and it was prepared by Thomas Brunner, Yu-Han Liu, and Moni Pande. The research focuses on a practical question: Are real-world attackers currently exploiting indirect prompt injection, and what outcomes are they seeking to achieve?

How Does Indirect Prompt Injection Work?

This technique differs from direct injection, in which a user attempts to bypass a chatbot’s restrictions through what is known as a jailbreak. In indirect injection, an artificial intelligence system processes external content, such as a web page, email message, or document, that contains malicious instructions. If the system treats these instructions as commands, it may execute them silently instead of adhering to the user’s original objective.

Google considers this threat an area of collaboration between Google DeepMind researchers and defenders at the Google Threat Intelligence Group. The risk is particularly significant as the use of artificial intelligence agents expands, enabling them to browse websites, process documents, and take actions on behalf of users.

The Common Crawl Survey and the Problem of False Positives

Google chose the public web as a relatively easy-to-monitor channel, since attackers can plant instructional text inside websites that they hope artificial intelligence systems will read. For accessibility and reproducibility, the company used Common Crawl, a large repository of web pages collected from the English-language web. The repository provides monthly snapshots containing approximately 2 to 3 billion pages.

These data consist primarily of static websites, including blogs, forums, comments, and self-published content. However, they do not include most social media platforms, such as LinkedIn, Facebook, and X, because Common Crawl skips sites that require login or use directives that prevent crawling. Google therefore said that studying prompt injection activity on social networks would be part of a separate research effort.

The survey faced a major problem involving false positives. Initial experiments found large quantities of harmless prompt injection text, particularly in research papers, educational materials, and cybersecurity articles discussing these attacks. To reduce ambiguity, Google followed a gradual filtering process:

  • Pattern matching: First searching for pages containing common indicators, such as phrases like “ignore the instructions” or “if you are an AI system.”
  • Classification using a large language model: Processing the candidate results with Gemini to determine the text’s intent and assess whether it was consistent with the document’s subject or appeared to be intrusive and suspicious.
  • Human verification: Conducting a final manual review of the classified results to increase confidence in the conclusions.

Google explained that this methodology is not comprehensive and may miss unusual patterns, but it provides a starting point for understanding the nature of prompt injection that actually exists on the web.

What Objectives Did Google Identify?

The cases found in the analysis ranged from harmless experiments to attempts with malicious objectives. Some website owners used hidden instructions to change the tone of the assistant reading the site, while others attempted to direct artificial intelligence summaries to add context they considered useful to readers. Google considers these cases harmless as long as they do not prevent summarization or seek to mislead the user, but they may become malicious if they include false information or direct the user to external websites.

The analysis also identified the use of prompt injection to improve visibility in search engines, with some websites attempting to push assistants to promote their businesses at the expense of other sites. Alongside the simple examples, more sophisticated attempts appeared that, according to Google, seemed to have been created by automated search engine optimization tools and then inserted into website text.

In other cases, instructions were used to prevent artificial intelligence agents from retrieving website content. The examples were not limited to a direct phrase such as “if you are an AI, do not crawl this site”; they also included attempts to lure the machine reader to a page that streamed an infinite amount of text and never finished loading, potentially wasting resources or causing timeout errors during processing.

As for the malicious cases, they included a small number of data theft attempts. However, Google said their level of sophistication was low, and it did not detect large volumes of advanced attacks, including some known data-extraction prompts published by security researchers in 2025. It also found websites attempting to sabotage a user’s device through the assistant, with some instructions seeking, if executed, to delete all files on the device. Google considered these examples unlikely to succeed because of their simplicity.

Upward Trend Despite Limited Sophistication

Google’s findings indicate that the observed activity does not prove that indirect prompt injection attacks have become a widespread production technique, as many cases appeared to have originated with individuals conducting experiments or pranks. However, the 32% increase in the malicious category between November 2025 and February 2026, after the survey was repeated across multiple versions of the archive, points to growing interest.

Google links the possibility of the threat escalating to changing cost-benefit calculations. In the past, these attacks were difficult to execute, and compromised systems did not reliably carry out malicious actions. Today, artificial intelligence systems have become more capable, increasing their value as targets, while attackers have begun using agentic artificial intelligence to automate their operations and reduce the cost of attacks. Accordingly, Google expects the volume and complexity of prompt injection attempts to increase in the near future.

In response, Google said it continues to strengthen its artificial intelligence models and products and uses red teams to test Gemini’s ability to resist adversarial manipulation. The artificial intelligence vulnerability rewards program also allows external researchers to participate in these efforts. The company emphasized that it will continue sharing information with the security community, while acknowledging that the current survey does not cover social platforms and does not provide a complete picture of all threat activity.

News source
Google Security Blog
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news