Introducing Layered Guardrails

Ⓐ Security's ironclad approach to safe Autonomous Offensive Security.

Offensive security guardrails shouldn't rely on the agent to police itself. Learn how Layered Guardrails keep autonomous agents safe without slowing them down.


The obvious response is to block anything uncertain. But an autonomous agent isn’t simply waiting for its next instruction. It’s pursuing an objective. Block a legitimate path often enough, and the guardrail itself can become an obstacle to overcome. This creates a paradox: guardrails that are too restrictive can encourage the very behavior they’re meant to prevent.

Simply telling an autonomous offensive agent to “stay in scope” isn’t enough. Its guardrails need to account for real-world context, enforce hard boundaries, and give agents enough freedom to be effective.

Our answer is Layered Guardrails: a new approach to offensive security guardrails that enforces safety across independent layers rather than relying on the agent to police itself.

How Layered Guardrails Work

In practice, safety instructions tend to break down for three reasons:

  1. Lack of situational awareness. An AI agent lacks implicit knowledge of the real-world context surrounding the systems it interacts with.
  2. Conflicting instructions. Agents optimize toward an objective. When achieving that objective conflicts with a safety instruction, the objective can win, even when doing so crosses a boundary the agent was instructed to respect.
  3. Context pollution. Over long runs with growing context and many instructions, safety constraints defined earlier in the context can lose influence.

Layered Guardrails address each of these failure points with independent layers of enforcement. No single control is trusted to keep the agent safe. Each layer plays a different role, so a failure in one can be caught by another before it becomes an unsafe action.

1. A Unified Classification System for Tasks, Actions, and Endpoints

We classify overarching tasks, individual actions, and target endpoints across four dimensions.

  • Operation type: What kind of action or task is being performed (e.g., read, write, modify, delete).
  • Data ownership: Whose data, endpoint, or environment is being touched (e.g., sandboxed test targets vs. live customer assets).
  • Reversibility: Whether the execution on a target endpoint can be undone if something goes wrong.
  • Real-world impact: The potential consequence of executing tasks on production systems or end-user endpoints.

Using these classifications, customers can select a guardrail preset or define custom policies, while certain safety boundaries (such as denial of service, destructive actions, or harm to users) cannot be disabled.

‍

2. Task-Level Preemptive Blocking

Checking agent actions only as they happen is often too late. Instead of relying on passive prompt instructions during execution, Layered Guardrails identify and block dangerous tasks before the agent starts.

Every exploitation lead our agents pursue is reviewed prior to running. If a task violates the safety policy during review, it’s blocked before the agent can commit to it. When a task is blocked, the work doesn’t stop; the agent simply replans the exploitation in a safer manner.

3. The Independent, Reasoning-Blind Judge  

During a long test run, safety rules buried in system prompts suffer from attention decay. By inserting an independent checkpoint into the process, Layered Guardrails don’t rely on the agent remembering its instructions.

Instead, state-changing actions are intercepted and evaluated by an independent, reasoning-blind judge. Reasoning-blind means the judge isn’t influenced by the agent’s reasoning. The judge reviews only the raw mechanics of the action, stripped entirely of internal reasoning or self-justifications, ensuring that attention decay or a long history of execution doesn’t dilute the evaluation.

4. Shared Memory Across Autonomous Agents

When multiple autonomous agents run in parallel without shared memory, they overlook key context. Layered Guardrails enable our agents to maintain a centralized memory that tracks every object created and changed during the engagement. This is critically important for actions where safety depends on hidden context (like whether an ID belongs to a test account or a real customer). If the object was created by the campaign, it's marked safe; otherwise, it's restricted.

This turns unanswerable runtime questions into a simple lookup.

5. The Autonomous Agent Safety Proxy

Relying solely on in-agent guardrails or model prompts creates a single point of failure. To overcome this, Layered Guardrails route agent traffic through an independent network proxy outside the agent’s control. Prohibited assets get hard proxy rules so the agent can’t reach them, regardless of what it "thinks."

When a block happens, the proxy responds with a structured error message explaining the rule hit and offering a safer path, preventing the agent from getting stuck in an endless retry loop.


Working together, these layers create several system-wide controls: agents can’t approve their own state-changing actions, blocked actions can’t be retried with different wording, and every intervention is logged for auditability. Read-only operations can proceed without the deeper evaluation reserved for state-changing requests.

Proving Guardrails Work

Building strong guardrails is difficult. Knowing whether they actually work is even harder. Security findings are relatively straightforward to benchmark: a vulnerability is real or it isn’t. Guardrails have to be evaluated against something that didn’t happen. Did the system prevent a dangerous action because the guardrail worked, or did the agent simply never attempt it? And just as importantly, did the guardrail block legitimate actions that should have been allowed?

To answer those questions, we built a tripwire application that records every dangerous operation attempted during an exploitation campaign, along with the agent responsible for it. We can run identical campaigns under different safety policies and compare what the agents attempted, what the guardrails blocked, and what they allowed through.

We test this using ARENA. ARENA generates fresh, realistic applications with known vulnerabilities and real-world protections. Because we know the ground truth (and because each target is newly generated), we can measure agent performance without relying on targets the model may have encountered before.

Instead of assuming a guardrail works because its rules look correct, we can observe how it behaves during actual autonomous exploitation.

Safe Offensive Security at Machine Speed

‍Guardrails impose restraints, but the same machinery that creates hard safety boundaries also lets autonomous agents move faster within them. The goal isn't to constrain autonomous offensive security. It's to give agents the freedom to operate at machine speed without sacrificing control.

To see Layered Guardrails work in practice, request a personalized demo today.

Share this article
About the author
Yarden Margalit has spent 10 years in offensive cybersecurity research. Before joining A Security as a founding engineer, she led research and offensive teams at Unit 8200, finishing as a group lead, building the kind of attacker capability that A Security now puts in defenders' hands.