WAF Bypass in the Agentic Era

What Happens When AI Adapts

See how autonomous AI agents adapt to WAF defenses, discover bypasses, and share what they learn, changing what web security looks like in the agentic era.

Traditional cybersecurity defenses weren’t built for today’s attackers. Agentic attackers now reason, adapt, and never tire. They read every rejection, work out why it was blocked, and rewrite its payload. They do this continuously and in parallel, at machine speed, and it never gets bored.

When one agent finds a way through, the rest pick up that knowledge right away.

This breaks a quiet assumption built into much of the security stack: attackers eventually run out of variations, or patience, and move on to easier targets.

Agentic attackers change that equation. A control that stops 99 out of 100 attempts looks very different when the attacker can generate the hundredth, learn from every failure, and keep going indefinitely.

The clearest example of this new paradigm is the Web Application Firewall (WAF).

Why WAFs work

It sits in front of a web application and inspects incoming HTTP/HTTPS traffic before it reaches the origin. It usually runs inline as a reverse proxy, though it can also work as a transparent bridge, a network appliance, a host-based module such as ModSecurity, or a cloud service at the edge.

It does several things very well.

It ships with managed rules and signatures, mostly regexes and pattern matchers, that recognize the shape of known attacks: SQL injection, cross-site scripting, command injection, path traversal, file inclusion. When a request matches, the WAF can block it, log it, or challenge the user. Before it matches anything, it normalizes the input to defeat obfuscation, decoding JavaScript escapes so ‘\u0041' turns back into A, stripping inline comments like /**/, collapsing odd whitespace, and unwinding command-line tricks like 'l'"s" into ls.

It also inspects the request itself, enforcing HTTP compliance, size and parameter limits, and allowed methods, so attacks that abuse the structure of a request get caught alongside the ones that hide in its content. And when a serious CVE drops, vendors push virtual patches that block the known exploit at the edge while the underlying software gets fixed.

A well-deployed WAF filters the constant background noise of automated scanners, raises the cost of an attack, and buys time during an incident. None of that is in question.

But notice what all of it quietly assumes: that a malicious request looks different from a legitimate one, and that whoever sends it will run out of variations before the WAF runs out of rules. That assumption held for a long time.

An agentic attacker is the thing that stops respecting it, and that is where the false sense of security begins.

Why WAFs provide a false sense of security

WAFs aren’t badly built. However, adaptive agents can do things WAFs weren’t meant to cover. It inspects HTTP between the client and the server, so some attacks stay out of view entirely. DOM-based XSS executes in the browser and may never reach the server at all. Authorization and business-logic flaws such as BOLA, IDOR, workflow abuse, and privilege escalation ride inside requests that look completely valid. A call to /api/invoices/123 is indistinguishable from a legitimate one, and only the application knows whether this particular user is allowed to read invoice 123.

Second, there are the tradeoffs any inline control has to make. To avoid breaking production traffic, a WAF balances coverage against false positives, parses the content types and encodings it can, and caps body size, parameter size, and transformation depth. Those are the right engineering calls, not weaknesses. Taken together, though, they leave a surface of shapes the WAF does not fully normalize or inspect.

For a human attacker, finding the one encoding, the one unparsed body, or the one normalization gap on that surface is slow, tedious work, and most attackers give up long before they get there. That giving up is the load-bearing assumption. Take it away and the false sense of security is exposed for what it is.

How AI agents bypass WAF

An agentic attacker does not treat that surface as a wall. It treats it as a search space, and it brings a few properties no human attacker has.

Every block is information. A rejection tells the agent something about the ruleset, so instead of moving on it forms a guess about which normalization step or signature caught the payload and rewrites against that guess. It does not spray random variants either. It narrows in on how the WAF actually behaves, which transformations defeat which tricks, which encoding the parser skips, where a size or depth limit leaves content uninspected. Many agents can do this at once, each working a different corner of the encoding and normalization space, which compresses days of manual fiddling into minutes.

The last property is the one with no human equivalent. When an agent lands a working bypass, it writes that result to a shared store, and every other agent can reuse it, against the same target or any target running a similar WAF configuration. One discovery becomes something the whole fleet knows. This is the collective-memory effect, and it is what turns a single lucky payload into a repeatable capability.

The WAF is doing its job correctly, and the wall of blocked requests on the dashboard is real. But those blocks are just the early, discarded attempts of a process that keeps adapting until it finds a payload the WAF was never set up to normalize. On that dashboard, "blocked" quietly stops meaning "safe."

How we tested this

We did not go looking for this in a lab first. We kept seeing it during real engagements, where our offensive security agents adapted past WAFs that were protecting live client applications. We are not going to put a customer's production system on display to make the point, so we reproduced the same behavior on our own ground.

That ground is “WorkSpend”, a deliberately vulnerable web application we built at A Security. Using ARENA, we can create realistic applications with known vulnerabilities specifically for testing autonomous offensive agents. We put WorkSpend behind a real Envoy WAF so our agents faced production-style edge filtering rather than a lab toy.

The attacker was our own autonomous exploitation agent, running as part of a multi-agent campaign. Each agent worked from an HTTP tool and a real browser, so it could see status codes, block pages, and full response bodies, and it fed every result back into its next attempt. The agents also shared a knowledge store that persists insights across runs and across agents, which is what lets one agent benefit from another's work.

We only counted a bypass when the underlying vulnerability actually fired, not when a request merely slipped past the WAF. For a cross-site scripting case that meant a unique marker value executing in a real, authenticated victim browser. For a server-side request forgery case that meant a live response coming back from an internal service and from the cloud metadata endpoint. Every confirmed bypass was filed as a finding and put through a separate reproduction step.

The two cases below come straight out of those runs.

What we’ve learned

Case one: the agent cracks the WAF on its own

Our first agent went after WorkSpend's login page, where a redirect parameter called next ends up in a dangerous browser sink. Before it touched the target it checked the shared knowledge store, the same store that does the heavy lifting in case two. This time the store came back empty. There was nothing to reuse, so the agent had to work it out for itself.

It started by pulling apart the app's own JavaScript bundle, and it found the flaw in the client code:

const j = next || "/dashboard";
return /^(https?:|\/\/|javascript:|data:)/i.test(j) && (window.location.href = j), j

The client only rejects a value that fails that test, so a javascript: URL passes the check and gets assigned straight to window.location.href. The agent's own note called it "inverted validation." That is a DOM-based cross-site scripting bug, and it lives in the browser, which is already outside what the WAF can see.

Then it hit the WAF. A plain javascript: scheme in next came back 403, and so did data:, in every case and encoding the agent tried:

javascript:console.log(1)	->	403
JavaScript:console.log(1)	->	403
data:text/html,<script>...	->	403
//evil.com	->	200

With no help from memory, the agent reasoned about the filter and ran a single sweep of 34 scheme-obfuscation variants against the endpoint. One family got through:

javascript:console.log(<marker>) -> 403 blocked
java%09script: console.log(<marker>)	-> 200 passed	(tab between "java" and "script")
java<TAB>script: console.log(<marker>)	-> 200 passed

A horizontal tab inside the scheme is all it took. The browser still reads java<TAB>script: as a valid javascript: URL, but the WAF signature does not match it. The agent then drove a real, authenticated victim browser to the crafted link and confirmed the payload actually executed rather than just reflected: the unique marker printed to the console, and an unauthenticated visit did not fire it.

Winning request:

GET /login?next=java%09script:console.log(<marker>)

This became the finding "DOM-based Cross-Site Scripting on Login Page via Next Redirect Parameter," rated high severity and confirmed by an independent judge.

The whole thing, from the first 403 to a browser-confirmed bypass, took about ninety seconds and a single fuzzing pass. No human attacker moves that fast, and most would have stopped at the first 403.

Case two: the agent skips the search entirely

Different campaign, a different corner of WorkSpend: a webhook test feature at POST/api/webhooks/test. You hand it a URL and the server makes a live request to it, then hands you back the full response. That is a textbook server-side request forgery sink, and the same Envoy WAF sits in front of it. The WAF blocks any request whose body contains an internal or cloud-metadata address like 169.254.169.254, returning a plain 403 Forbidden page. Importantly, it matches literal strings in the body, not destinations.

An autonomous system should not start every campaign from zero. That’s why a powerful harness, like the one we’ve built at Ⓐ, is so critical. The harness around our agents manages context and turns useful discoveries into persistent memory that can be shared across agents and across runs. As the system tests more applications, that body of knowledge grows with it.

This time the agent did not fuzz from scratch. After the first 403 it queried the shared knowledge store with a very specific question:

  • "WAF bypasses SSRF for WorkSpend: what encodings, IP formats, or payload transforms bypass the WAF when posting internal or metadata URLs like 169.254.169.254 in the JSON body to /api/webhooks/test?"

Collective memory answered with one confirmed bypass and one untested lead:

  • Confirmed: http://127.1:3000/ reaches the internal service, while http://127.0.0.1:3000/ is blocked. Same destination, different verdict, because the filter matches literal strings.
  • Untested hint: full-body character-set encodings such as UTF-16 were flagged as worth trying on this sink but not yet confirmed.

The agent reused the confirmed loopback bypass immediately and reached the app's own internal service:

POST /api/webhooks/test
{"url":"http://127.1:3000/","method":"POST","headers":{},"event":"expense.approved "}

response: 200, x-powered-by: Express, body: Cannot POST /

Then it took the untested hint and proved it. It re-encoded the request body as UTF-16LE, so the WAF's signature engine never saw the string 169.254.169.254, while the application still parsed the JSON and made the request:

POST /api/webhooks/test
Content-Type: application/json; charset=utf-16le (body, as UTF-16LE)
{"url":"http://169.254.169.254/latest/api/token","method":"PUT",
"headers":{"X-aws-ec2-metadata-token-ttl-seconds":"21600"}}

response: 200 server: EC2ws
body: <IMDSv2 session token, redacted>

The same request in ordinary UTF-8 returns a 403 from the WAF. server: EC2ws plus a live IMDSv2 session token means the forged request genuinely reached the AWS metadata service. The agent spent one lookup and a couple of requests where a human would have spent an afternoon. It did not invent the bypass. It read it, reused it, and confirmed the part the notes were unsure about.

This became the finding "Server-Side Request Forgery on Webhook Test Endpoint Reaches Internal Network and AWS IMDS," rated high severity.

Two halves of the same system

Put the two cases side by side and you can see the whole. In the first case the shared store was empty, so the agent had to crack the WAF itself. In the second, the store already held a bypass, so the agent skipped the search and spent its time confirming a hint instead. What one agent works out the hard way becomes what the next agent reuses in a single lookup. The shared knowledge store is not a static cheat sheet. It is a living record that agents both write to and read from, and that is a property a lone human attacker, working from patience alone, simply does not have.

What this means for defenders

  • Treat the WAF as a speed bump, not a wall. It filters noise and buys time, but it does not remove the vulnerability, so a green WAF dashboard should never be the thing that closes a ticket.
  • Stop assuming attacker patience runs out. Rate limiting and "they will get bored and leave" no longer hold, and any control that leans on them needs a second look.
  • Move toward controls that understand intent. Authorization checks, application context, and behavioral detection catch what payload-shape matching cannot, because an agent cannot re-encode its way around a permission it does not have.
  • Watch for machine-speed adaptation. A flood of systematic payload mutations against one endpoint is itself a signal. People do not probe like that.
  • Borrow the attacker's best trick. Shared memory cuts both ways, and defenders who pool what they learn at machine speed close the same gap the attacker is exploiting.

Where this leaves us

A WAF is worth having, and it is trusted for good reason. It filters noise, raises the cost of an attack, and buys time during an incident. But that same trust is the risk. A WAF defends against a shape of attack, and an agentic attacker changes shape faster than any signature set can keep up with. The techniques themselves are old. What is new is an adversary that reads every rejection, reasons about the ruleset, keeps going without getting tired, and shares whatever it learns with every other agent in the fleet.

The shift this forces is away from "block the known-bad shape" and toward controls that understand intent, authorization, and application context.

Interested in learning more about how A Security’s offensive agents adapt to WAF defenses and discover bypasses? Request a demo today.

Share this article
About the author
Tomer Zait has spent his career across application security, working as a WAF integrator, penetration tester, application security engineer, and security researcher. Before joining A Security as Security Research, he was a Principal Security Researcher at F5, presented his open-source tools at Black Hat Arsenal, and won seven Israeli CTF competitions.