Cloudflare tested its web application firewall (WAF) using advanced AI models capable of modifying attack requests after each attempt, rather than relying solely on a fixed set of tests. The company conducted the experiment in a customer-specific staging environment with prior approval and recorded 1,107 attempts across 45 scenarios.
The core idea was to simulate part of an attacker’s behavior: trying different encodings, moving the payload to other locations within an HTTP request, or changing how the destination was represented, then using the response to select the next attempt. The model was not given the application’s source code, the internal WAF rules, or rule numbers; it dealt only with the request context and specific HTTP response data.
How Did the Adaptive Test Work?
Each scenario began with a request known to be blocked by the WAF, after which the model proposed a new modification. After the request was sent, another model or a review call from the model examined the response before the system determined the next step within a maximum attempt limit. The code handled request execution and evidence logging. Before each attempt, it also verified the domain name against an allowlist, stopped redirects, and prevented the system from changing or deploying protection rules.
Forty-four of the scenarios covered six categories: cross-site scripting (XSS), SQL injection, command injection, server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j attacks. The forty-fifth scenario addressed log injection separately.
From 1,107 Attempts to 49 Actionable Findings
The WAF performed strongly overall, with near-complete coverage of the XSS, LFI, SQLi, and Log4j categories. After human review, 49 findings remained worthy of investigation, 48 of which belonged to the command injection and SSRF categories. A total of 558 requests were blocked, while other attempts were excluded because they were invalid, benign, duplicated, or did not reach the target.
Cloudflare did not consider an unblocked request a confirmed vulnerability. It required the request to be valid and to reach the target, to remain clearly malicious, to be attributable to the WAF, and to be safely reproducible by engineers. This distinction is important because bypassing a WAF alone does not prove that the application was successfully exploited or that sensitive data was accessed.
An Example of a Gap in SSRF Detection
In one scenario, the system changed how the address of the cloud metadata service was written, used different numeric representations, and placed the address in multiple parts of the request. Most attempts were blocked, but one attempt using a trailing-dot representation resulted in a redirect instead of being blocked by the WAF. Cloudflare considered this a signal worthy of investigation, not evidence of access to credentials or successful exploitation.
What Changed in Practice?
Cloudflare reviewed the findings to determine whether it needed a new rule, improved request normalization, or intervention from another security layer. The results contributed to three changes in the Managed Ruleset: adding SSRF - Obfuscated Host and SSRF - Restricted Protocol detections in the July 21 release, and improving SSRF - Cloud detection. The obfuscated-host detection came directly from requests that used unusual numeric formats for internal addresses.
The experiment shows that AI here is a tool for expanding the testing space, not a substitute for engineering judgment. Two models from the same family produced different variations, but the fundamental issues appeared in both. Increasing the number of attempts within a single scenario also did not guarantee the discovery of additional findings; some sequences began repeating ideas, while increasing the number of starting points, attack categories, and input locations provided broader coverage.
WAF limitations remain: a payload that bypasses the firewall succeeds only if the application itself is exploitable, so updating software and fixing vulnerabilities remain necessary. Cloudflare also recommends first enabling rules in logging mode, reviewing matching requests and security events, and then moving to blocking after confirming that legitimate traffic is not affected. The company later plans to present a white-box test in which the model knows both the application vulnerabilities and the WAF rules.