Back to Newsroom
Threat Intel

Google's AI Agent Found 500+ XSS Bugs — And Proved They're Exploitable

An internal Google AI agent didn't just flag XSS risks — it built working exploit chains to prove them, slashing false positives and surfacing a cache-poisoning bug with region-wide reach.

Google's AI Agent Found 500+ XSS Bugs — And Proved They're Exploitable

Google has disclosed results from an internal AI-driven security agent that spent months hunting cross-site scripting (XSS) vulnerabilities across its own web applications — and, in a notable departure from typical automated scanning, didn't stop at flagging suspicious code. The system builds and runs proof-of-concept exploit chains before anything reaches a human engineer.

What happened

The agent, built primarily on Google's Gemini models, began as an internal effort in late 2025 and expanded into a full program in early 2026. Rather than relying on pattern-matching to flag "risky-looking" code — the approach responsible for much of the noise that plagues automated scanners — it pairs an AI-driven triage step with a purpose-built validator. For XSS specifically, that validator injects JavaScript into a browser-like test environment and confirms the payload actually executes before a finding is ever surfaced.

The result, according to Google, is a near-zero false-positive rate across tests that also cover SQL injection, path traversal, remote code execution, and server-side request forgery. Of several hundred applications built on Google's hardened internal web frameworks, the agent surfaced only two confirmed XSS issues — both in internal tooling or debug endpoints with weaker guardrails than production services, reinforcing the value of consistent framework-level controls.

Exploit chains, not just findings

The more serious discoveries weren't single bad inputs — they were multi-step chains:

  • Cache poisoning via an unvalidated path segment. A JavaScript-serving endpoint accepted a URL path segment that was reflected into the response but excluded from the cache key. That gap meant a malicious response could be cached and served to other visitors in the same region. Google found no evidence of in-the-wild abuse, but the chain could reach sensitive Google-operated domains and any third-party site loading the affected script.
  • A signed-redirect bypass on an admin console. An unvalidated redirect parameter could reach window.location, but a cryptographic signature requirement initially blocked exploitation — until the agent located a separate authorization endpoint that would issue a valid signature for an attacker-controlled JavaScript URL, turning a "protected" redirect into a working XSS primitive.
  • A browser-extension nonce/postMessage weakness. Loose checks on an extension's external-connection allowlist, combined with a recoverable one-time nonce and permissive message forwarding, let attacker-controlled content reach a page under active debugging. Because the flow also accepted data: URLs, this escalated into arbitrary JavaScript execution — a universal XSS condition.

Why it matters

This isn't just a story about Google patching bugs in its own products — it's a signal about where automated vulnerability research is heading. Pairing an AI triage layer with deterministic, execution-based validation is a meaningfully different approach from scanners that just pattern-match and hand analysts a pile of maybes. The cache-poisoning and signed-redirect chains are good illustrations of how individually low-severity gaps — an unvalidated path segment, a redirect parameter — combine into high-impact, browser-side code execution.

Google notes its own validators can still miss genuine issues, and that human engineers continue to review every proposed fix before it ships — a reminder that AI-assisted discovery is a force multiplier for triage, not a replacement for review.

What to do

  • Audit cache-key construction on any endpoint that reflects path segments, query parameters, or headers into a cached response — if the reflected value isn't part of the cache key, you likely have a poisoning vector.
  • Re-check redirect validation end-to-end, including any signing or authorization service that issues tokens for redirect targets — a signature check is only as strong as every endpoint that can produce a valid signature.
  • Review browser-extension message handling: validate the sender origin on every postMessage listener, treat one-time nonces as single-use server-side (not just client-side), and never allow data: URLs where a script-execution context is reachable.
  • Combine AI-assisted scanning with execution-based validation in your own pipeline — flagging risky patterns is cheap; proving exploitability is what actually cuts triage time.
SHARE